Your creative fieldWe are looking for a full-time, permanent Senior Platform & Reliability Engineer (all genders) to start as soon as possible.
We live remote-first, but you have the freedom to choose whether you want to work hybrid or completely on-site due to your proximity to one of our locations (Berlin, Cologne, Hamburg, Munich).
As a Senior Platform & Reliability Engineer at Contabo, you take architectural ownership of our shared infrastructure services – the foundation that multiple development teams build on: API gateway and ingress (Kong, Nginx Ingress), persistent storage (Ceph, Longhorn), secrets and identity infrastructure (Vault, Keycloak), edge security and WAF (Cloudflare), and our observability stack.
You'll be taking over grown, partially under-documented systems – and that's exactly where the appeal of this role lies: you work your way deep into these systems, identify and remediate known weak points, evaluate aging components with solution-agnostic build-vs-buy reasoning, and decide which legacy pieces get fixed, replaced, or retired.
One of your central mandates: you design and establish an on-call process, including runbooks for platform and infrastructure incidents – where no formal process exists today.
In parallel, you mature our observability practices, rolling out distributed tracing, SLOs/SLIs, and meaningful dashboards across the stack, building on our existing tooling with Prometheus, Grafana, Alloy, and OpenTelemetry.
In system design reviews, you bring your strong grounding in the fundamentals – load balancing, caching, sharding, replication, consistency trade-offs – and apply them to new and existing services alike.
You won't be managing a team, but you will be the technical authority multiple teams rely on: you advise across teams on shared infrastructure services, document your knowledge consistently, and actively distribute it – so that critical know-how never again depends on a single person.
Success in this role means: known risk.