Remote job
Senior Site Reliability Engineer (SRE)
Job details
About this role
Role overview A senior Site Reliability Engineer is needed to help establish and grow a high-impact SRE function within a cloud-native contact center platform. The role acts as a technical leader shaping how reliability is designed, measured, and improved across critical customer-facing services, with a strong emphasis on automation, observability, and cross-team influence.
Responsibilities - Drive improvements to reliability, scalability, and performance across production services handling customer interactions. - Define SLIs, SLOs, and error budgets, and use them to inform engineering priorities and trade-off decisions. - Build and evolve observability systems covering metrics, logging, tracing, and alerting so that on-call engineers get actionable signals. - Lead complex incident response, including serving as incident commander when needed, and run blameless postmortems that drive systemic corrective actions. - Reduce operational toil through automation, tooling, and improved workflows, and partner with product and platform teams on architecture and production-readiness decisions. - Mentor other engineers, raise the organization's operational maturity, and create paved-road tooling that makes it easier for teams to operate services reliably.
Requirements - 6–10+ years of experience in SRE, infrastructure, or backend systems engineering. - Proven track record of owning reliability outcomes for complex, distributed production systems. - Strong hands-on experience with at least one major cloud provider (AWS, GCP, or Azure) at production scale. - Deep understanding of observability, incident management, and system performance tuning. - Proficiency in at least one programming language such as Go, Python, or Java, with a focus on automation and tooling. - Ability to influence how other teams work without direct managerial authority, and sound judgment during high-pressure incidents.
Nice to have - Prior experience building or scaling SRE practices, including SLO frameworks, incident processes, and on-call models. - Kubernetes or other container orchestration experience and Infrastructure as Code using Terraform or similar tools. - Background in performance engineering, capacity planning, or working on high-growth, rapidly scaling systems.
Benefits and work setup - Annual US hiring range of approximately $140,000–$180,000, with placement depending on location, experience, education, and skill level. - Benefits package includes medical, dental, vision, a 401(k) plan, and wellness benefits. - Work authorization in the country of hire is required; visa sponsorship is not offered for this role.