Volver

Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

CompraTica Empleos

EMP:Technology
Berlin
Tiempo Completo
Remoto
0 vistas

Descripción

Introduction At a glance Location &workmodel:Berlin, hybrid Tech stack:Kubernetes on our own servers, Harvester (KubeVirt), Argo CD/Flux, Prometheus/Grafana, Longhorn/Ceph Team:A growing SRE team – you report to our CTPO for now and to the Team Lead SREwe'rehiring next; two system administrators in Pforzheim run the physical hardware Process:Intro call · take-home task (~2h) · 90-min tech interview with our developers · leadership conversation · meet the team Languages:Fluent Englishrequired; German is a plus, nota must Why this role is special Most SRE jobs today mean clicking around a managed cloud console.

This one doesn't.

We run our own hardware in Frankfurt and are building a modern private cloud platform on Kubernetes and Harvester – on-prem by default, with elastic burst into the public cloud and the option to go cloud-only later.

You won't inherit a finished SRE practice: you'll help define it, side by side with our Berlin development teams – and you won't do it alone, a Team Lead SRE hire is coming next.

SRE here is an enabling discipline: you build what our developers need to ship reliably, while two system administrators in Pforzheim run the physical hardware.

And the impact is direct – our product discovery technology powers more than 2,000 European online shops (Intersport, SPAR, Douglas and more), handling billions of shopper queries a year.

When product discovery is slow or down, our customers lose revenue in real time.

Your first 90 days You get to know both products, join the on-call rotation with a buddy, and own your first reliability topic – SLOs for one product, alerting that actually helps at 3 a.

, or automating away a piece of toil.

By day 90 you've shipped visible improvements and know where you want to take the platform next.

Your mission Define and own SLOs, SLIs and error budgets; drive data-informed reliability decisions Lead incident response end-to-end: fast detection, clear communication, blameless postmortems – and reduce whole.

¿Te interesa? Aplicá ahora