Remote job
Senior Site Reliability Engineer
Job details
About this role
Role overview The Senior Site Reliability Engineer owns the reliability, performance, and resilience of cloud infrastructure powering customer-facing products and AI/ML workloads. Sitting on a Platform Engineering team, this is an automation-first role focused on defining service-level objectives, leading incident response, and converting manual operations into monitored, hands-free processes. Because the underlying systems influence health outcomes for millions of people, maintaining the highest production-quality and security standards is central to the work.
Responsibilities - Own end-to-end reliability, performance, and resilience of AWS and Kubernetes environments, including those supporting AI/ML workloads, and define, measure, and uphold SLOs across critical services - Participate in an on-call rotation, lead incident response, drive deep-dive root cause analysis, and rigorously review infrastructure changes - Build and maintain monitoring, alerting, and observability systems that surface issues before users experience them - Translate ambiguous, high-performance scaling requirements into automated, composable infrastructure-as-code deliverables using Terraform, and proactively identify cost-efficiency and performance gains - Pay down impactful tech debt and reduce operational toil by applying AI tools and automation to convert repetitive work into hands-free, monitored processes - Establish deployment and observability standards that empower the broader engineering team to ship features faster and more reliably, and communicate complex reliability concepts to varied audiences - Ensure infrastructure and operations meet security and healthcare compliance obligations
Requirements - 4+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role - Deep expertise with Kubernetes and Terraform in a cloud-first environment, AWS preferred - Strong track record defining SLOs, building monitoring and alerting systems, leading incident response, and running blameless post-incident reviews - Solid software engineering fundamentals in Python or Go, applied to infrastructure automation, with Kubernetes API experience as a plus - Demonstrated experience driving cloud cost-efficiency and performance optimization across compute, storage, and networking - Experience supporting AI/ML or data-intensive production workloads is a plus - Experience operating in security-conscious or regulated environments such as HIPAA or SOC 2 is a plus - Fluency with AI-assisted engineering and operations workflows, or strong motivation to build it quickly
Nice to have - Background supporting AI/ML or data-intensive production systems - Familiarity with Istio, Datadog, GitLab, Postgres, NATS, or similar tooling - Comfort working in a fast-paced, mission-driven environment with high individual accountability
Benefits and work setup - Target base compensation range of $191,000 to $226,000, with eligibility for equity incentive and competitive benefits plans - Benefits include flexible paid time off, medical/dental/vision plan options, 401(k) with company match, flexible spending accounts, and a telehealth benefit - Remote-eligible position with occasional travel expected to the company's headquarters - Visa sponsorship is not available for this role at this time