Remote job
Site Reliability Engineer
Job details
About this role
Role overview Take responsibility for the reliability, performance, capacity, and security of a production platform. You will own shared infrastructure and cross-service operational work, while product engineers operate and secure their individual services. An AI-based alert triage system handles initial investigation; you supervise its diagnoses and improve its ability to reduce future incidents.
Responsibilities - Maintain production uptime, latency, and capacity, and lead incident response when needed. - Own the platform and cluster infrastructure, including compute, data stores, messaging, and observability systems. - Improve paging, escalation, and alert quality to reduce unnecessary interruptions. - Set security standards and tooling, and handle platform-level and cross-service security work. - Manage cloud security posture, infrastructure-as-code reconciliation, GitOps, and the security backlog. - Deploy and assess the AI first-responder, validating its diagnoses and expanding the work it can safely handle.
Requirements - At least four years of relevant experience. - Ability to take end-to-end ownership of production reliability and platform operations. - Experience with on-call incident response and improving operational systems. - Capability to oversee infrastructure, observability, and security across shared services. - Willingness to supervise automated incident triage and use its findings to improve response processes.
Benefits and work setup - Full-time remote role with primary on-call ownership included in the position and base salary. - Work is organized around outcomes rather than set hours; recovery time after overnight incidents is expected, and planned time off is covered by a backup.