Remote job
Site Reliability Engineer
Job details
About this role
Role overview A Site Reliability Engineer role centered on running client-facing platforms where reliability is a business-critical promise. The work combines SLO ownership, proactive failure rehearsal, and aggressive automation to remove repetitive operational toil. The position is fully remote across EMEA and suited to engineers who stay calm when systems don't.
Responsibilities - Define, monitor, and defend service-level objectives and error budgets for production platforms. - Rehearse failure scenarios through game days, chaos exercises, and structured drills. - Lead incident response, acting as incident commander during high-severity events. - Automate recurring operational tasks to progressively reduce manual intervention. - Operate Kubernetes-based workloads, Terraform-managed infrastructure, and observability tooling as day-to-day instruments.
Requirements - Track record of production ownership for systems governed by real SLAs. - Practical experience writing and defending error-budget policies. - Daily fluency with Kubernetes, Terraform, and modern observability stacks. - Incident command experience, with a calm and structured approach under pressure.
Benefits and work setup - Compensation benchmarked annually against the external market. - Remote work from any major hub city, with relocation support. - Hardware and learning budgets available without gating approval. - Minimum six weeks of paid leave, modeled visibly by leadership.