Remote job
Site Reliability Engineer
Job details
About this role
Role overview Join a small infrastructure team as a Site Reliability Engineer responsible for keeping a production platform dependable for external customers. The role centers on owning the on-call rotation, the incident-response muscle, and the systems that surface how the platform is really behaving. It is a remote, full-time position with a compensation range of $150,000 to $400,000 USD.
Responsibilities - Run the on-call rotation and serve as the primary responder for production incidents. - Build tooling, automation, and runbooks that reduce pain for the engineers who handle the next page. - Lead blameless post-mortems that convert outage learnings into concrete engineering work items. - Define, measure, and maintain SLOs that customers can rely on and that map to real reliability work. - Improve observability coverage and the signals the team uses to detect and triage issues early. - Partner with product and platform engineering to set capacity plans that match growth.
Requirements - Production SRE experience at a company with public-facing SLAs. - Strong, well-formed opinions about observability, error budgets, and sustainable on-call practices. - A track record of treating outages as preventable, supported by specific examples. - Comfort writing code to build internal tools and automate operational toil. - Ability to lead incident response calmly and communicate clearly under pressure.
Benefits and work setup - Remote-friendly full-time role. - Hiring process includes a take-home coding screen, two technical interviews, and a 30-minute conversation with the founder, with an offer target of roughly five business days after the final interview.
Nice to have - Prior ownership of capacity planning at meaningful scale. - Experience shaping SLO programs from a blank page rather than inheriting one.