Remote job
Senior Site Reliability Engineer
Job details
About this role
Role overview
A senior engineering opportunity focused on building and scaling cloud infrastructure for a healthcare platform serving patients with joint, back, and muscle conditions. The role sits within a newly formed Infrastructure & SRE group and involves close partnership with product engineering teams rather than traditional ticket-driven operations. It emphasizes hands-on technical execution, paved-road tooling for developers, and bringing automated discipline to production environments.
Responsibilities
- Standardize the AWS footprint with Terraform, eliminate manual provisioning, and integrate AI-assisted tooling for code review and infrastructure operations - Build universal CI/CD pipelines (GitHub Actions) and track DORA metrics to improve delivery velocity and reliability across engineering - Implement observability across AWS and third-party integrations, maintaining SLOs, SLIs, and error budgets for web, mobile, telephony, and data processing platforms - Act as a bridge between SRE and software or data engineering, building self-service tooling so developers can ship without filing tickets - Contribute to standing up an on-call rotation and escalation paths covering US time zones - Lead production incident response, run blameless post-incident reviews, and address systemic root causes - Maintain infrastructure controls such as IAM, encryption, and network segmentation aligned with HIPAA and HITRUST requirements
Requirements
- 5+ years in Software Engineering, SRE, DevOps, or Platform Engineering at a senior level, with a track record of delivering technical initiatives - Deep hands-on AWS expertise (VPC, IAM, ECS/EKS, Lambda, RDS, S3) and production-grade Terraform at scale including modules, state management, and multi-environment setups - Strong programming skills in Python, Go, TypeScript, or similar, treating infrastructure as software - Hands-on experience with observability stacks such as Datadog, CloudWatch, or Grafana, plus practical SLO and SLI implementation - Experience maintaining CI/CD pipelines and tracking delivery metrics like DORA - Willingness to travel up to 10% for onsite meetings, team collaboration, and events
Nice to have
- Experience operating in HIPAA, HITRUST, SOC 2 Type II, or comparably regulated growth-stage environments
Benefits and work setup
Remote-first with hybrid options available in Nashville, paid parental leave, generous PTO, medical, dental, vision, life, and disability insurance from day one with an HSA contribution, 401k with employer match, wellness resources, and an inclusive culture.