Remote job
Senior Site Reliability Engineer (Systems Engineer III)
Job details
About this role
Role overview A senior infrastructure-focused engineering role dedicated to owning the reliability, availability, and performance of a Google Cloud Platform stack that powers a telehealth product. The position partners closely with fullstack engineers to elevate DevOps maturity, embed AI-assisted tooling into operational workflows, and shape how the technology organization approaches incident response and infrastructure changes.
Responsibilities - Lead the reliability program end to end, including defining and operationalizing SLOs and SLIs, addressing known pain points such as endpoint latency and Cloud Run cold starts, and owning the disaster recovery posture through backup and restore testing, recovery objectives, and multi-region resilience work. - Build out observability and incident response on GCP Cloud Operations, transition alerting into a dedicated paging tool with clear severity categorization and escalation paths, and participate in the on-call rotation alongside other engineers. - Grow infrastructure-as-code and delivery practices by expanding Terraform coverage across the GCP footprint and hardening the GitHub Actions CI/CD pipeline so frequent releases stay safe. - Partner with IT leadership on security and compliance hardening, including IAM, secrets management, audit logging, and the infrastructure-side controls required for HIPAA and SOC 2. - Multiply the team's impact through documentation, pairing, code review, and bringing reliability input into new features during design. - Use AI tools fluently in day-to-day engineering work and help build AI-assisted operational tooling such as automated level-one triage and runbook helpers.
Requirements - Several years of production experience operating systems on Google Cloud Platform, including Cloud Operations and Cloud Run, and a track record of improving SLOs, incident handling, and on-call maturity. - Strong proficiency in TypeScript and Node.js, with the ability to navigate a fullstack codebase built on React and Firebase services such as Firestore, Authentication, and Hosting. - Hands-on experience with Terraform-driven infrastructure as code and CI/CD pipelines such as GitHub Actions. - Demonstrated ability to harden infrastructure for security and compliance, covering IAM, secrets, and audit logging. - Clear written communication and a habit of producing documentation such as runbooks, postmortems, and design docs, plus a pairing-oriented approach to keeping knowledge distributed. - Daily, hands-on use of AI coding and operations tools, with experience building automation on top of them rather than only prompting them.
Nice to have - Background in a regulated industry, ideally healthcare or telehealth, with hands-on experience implementing HIPAA or SOC 2 infrastructure controls. - Familiarity with healthcare systems, health outcomes, and benchmarks. - Personal motivation to expand access to evidence-based care.
Benefits and work setup - Remote-friendly position with a work-from-home stipend. - Medical, vision, and health insurance plans, with many options fully covered for employees. - 20 days (4 weeks) of paid time off, scaling to 5 weeks after two years and 6 weeks after five years, plus 10 company holidays. - 401(k) contribution platform and additional offerings such as life insurance, disability coverage, and virtual primary care through the benefits provider. - Compensation range of $140,000–$150,000 USD, with a transparent first-offer pay approach and opportunities for annual increases tied to company performance.