Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance.
As our GPU fleet continues to grow, we're investing in the infrastructure that operates it.
Our team builds the tooling, automation, and intelligent systems that make GPU infrastructure scalable, observable, and increasingly autonomous.
We work across production engineering, distributed systems, capacity management, and AI-powered operational tooling to help researchers and product teams move faster while maximizing the efficiency of every GPU.
This is a unique opportunity to work on infrastructure at the frontier of AI, where small improvements in fleet efficiency, reliability, and automation have an outsized impact on the development and deployment of AGI.
About the RoleWe're looking for a Software Engineer with deep experience operating large-scale GPU or compute infrastructure.
You'll design and build the systems that manage GPU clusters at scale—from fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention.
You'll partner closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and overall compute utilization.
This role is ideal for engineers who enjoy solving complex operational challenges, building internal platforms, and working on infrastructure that directly powers frontier AI.
In This Role, You WillDesign, build, and operate software that manages large-scale GPU infrastructure supporting ChatGPT inference.
Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead.
Improve observability, reliability, and operational efficiency across thousands of GPUs.