Our engineering team is distributed across Europe and North America, and we're looking for a Senior Site Reliability Engineer to take ownership of our cloud infrastructure and elevate our DevOps and reliability practices.
In this role, you'll evolve our Google Cloud Platform (GCP) infrastructure, mature our observability platform, drive incident management processes, and partner closely with Product & Engineering teams to ship reliable, high-quality software.
You'll be a key voice in championing SLOs, error budgets, and DORA metrics across the organisation.
What You'll DoDesign, evolve, and scale our cloud infrastructure on GCP.
Build tooling and automation that promote team autonomy and reduce toil.
Advance our observability platform, improving mean time to recovery (MTTR) and system visibility.
Build transparency into infrastructure costs and drive cost optimisation initiatives.
Champion reliability best practices including SLOs/SLIs, error budgets, and post-incident reviews.
Help engineering teams leverage GCP effectively and govern usage at scale.
What We're Looking ForRequired:3+ years of Site Reliability Engineering or production SRE experience.
Strong proficiency with Google Cloud Platform (GCP), including cost optimisation and governance.
Nice to Have / Additional.