Remote job
Senior Site Reliability Engineer
Job details
About this role
Role overview
Own the reliability, performance, and operational readiness of high-throughput, low-latency production systems. This is a hands-on senior engineering role spanning observability, infrastructure, incident response, capacity planning, safe delivery, and developer tooling, with direct accountability for how critical services behave under real customer traffic.
Responsibilities
- Define and operate SLIs, SLOs, error budgets, dashboards, and actionable alerts for critical production paths. - Lead investigations and recovery during major incidents, then turn findings into practical postmortem actions. - Design secure, resilient, and cost-conscious infrastructure with appropriate timeouts, retries, backpressure, load shedding, graceful degradation, and blast-radius controls. - Perform load testing, profiling, saturation analysis, and capacity planning ahead of expected growth. - Improve deployment safety through progressive delivery, automated rollback, stronger pre-production signals, and infrastructure-as-code practices. - Build operational software, runbooks, failure-testing exercises, and production-readiness processes that reduce toil for engineering teams.
Requirements
- 6–10 years of experience in SRE, production engineering, infrastructure, or backend engineering in primarily cloud-based environments; AWS experience is preferred. - Demonstrated ownership of a significant system through design, delivery, production operation, and remediation of failures. - Practical experience defining and applying SLIs, SLOs, and error budgets. - Experience serving as a lead or primary responder for serious, customer-facing incidents and improving organizational learning afterward. - Strong understanding of distributed-system failure modes, including saturation, cascading failures, retry storms, capacity limits, degradation, and load shedding. - Hands-on knowledge of cloud networking, load balancing, containerization, Kubernetes/EKS, Terraform or equivalent, and production programming in Go, Python, or a comparable language.
Benefits and work setup
- Fully remote work model, subject to authorization to work from the chosen home location and applicable country restrictions. - Participation in an on-call rotation, with an emphasis on improving runbooks, escalation clarity, and pager health. - An inclusive environment that welcomes varied backgrounds and perspectives.