You will work across our cloud infrastructure (AWS/GCP), Kubernetes (EKS), CI/CD pipelines, observability, and platform security to improve release confidence and operational excellence.
In this role, you will partner closely with engineers across backend, frontend, data, and AI teams to remove operational bottlenecks, improve developer workflows, and standardize how services are built, deployed, and operated.
This is a hands-on role with clear impact, directly contributing to outcomes such as more reliable CI pipelines, faster incident detection, and improved infrastructure cost visibility.
What you'll doImprove CI reliability & developer productivity: Reduce flaky pipelines and shorten feedback loops across CI/CD systems to improve developer experience and delivery speedStrengthen observability & incident readiness: Build actionable dashboards, alerts, and SLOs for key user journeys to help teams detect, diagnose, and resolve issues fasterOperate and evolve our Kubernetes platform: Manage and improve our AWS EKS environment, ensuring safe deployments, runtime reliability, and scalable infrastructureDrive infrastructure automation: Expand infrastructure-as-code practices using Pulumi and Type.