Your missionYou will make Lyceum's AI inference platform reliable, secure, and scalable - ensuring it performs under pressure as we grow to thousands of concurrent users.
While others on the team expand what the platform can do, your job is to make sure it keeps working, fails gracefully, and gets faster over time.
Your focusScalability: Architect and implement the systems that allow our inference platform to scale to thousands of concurrent users.
This includes request routing, load balancing, autoscaling, and resource scheduling across GPU clusters.
Reliability and observability: Build robust monitoring, alerting, and incident response tooling.
Design for graceful degradation, automatic recovery, and minimal downtime.
Identify and eliminate bottlenecks at every layer.
Infrastructure evolution: Evaluate and integrate open-source inference frameworks and tooling (Dynamo, vLLM, Triton, etc.
) where they improve throughput, latency, or stability of the serving stack.
Your KPIsPlatform uptime and availability (SLA adherence)P50/P95/P99 latency and throughput under loadTime-to-detection and time-to-resolution for incidentsScalability milestones (concurrent users, requests per second, GPU utilisation)Your profileWe consider candidates from diverse backgrounds, with a deep love for technical challenges and the desire to take on ownership beyond what's reasonably expected.
Requirements3+ years of experience in backend, infrastructure, or systems engineeringStrong proficiency in Go and PythonExperience building or operating a model serving platform or ML platformSolid understanding of systems performance - profiling, benchmarking, and optimising latency and throughputFamiliarity with observability tooling (Prometheus, Grafana, OpenTelemetry, or similar)Understanding of security fundamentals - network isolation, authentication, encryption, mu.