Remote job
Principal Site Reliability Engineer
Job details
About this role
Role overview
This is a senior technical leadership position shaping the long-term reliability, scalability, and operational excellence strategy for a large cloud-native platform. The role combines deep hands-on engineering with broad organizational influence, partnering across application, platform, security, and infrastructure teams to build resilient systems and modernize architecture. It also carries significant responsibility for mentoring senior engineers and aligning with global SRE counterparts on shared practices.
Responsibilities
- Define and drive the multi-year reliability strategy, setting technical direction and operational standards across services and teams. - Lead architecture decisions for highly available, fault-tolerant distributed systems running on AWS, GCP, Kubernetes, and GKE. - Establish and mature SRE practices including SLIs, SLOs, error budgets, production readiness reviews, service ownership standards, and operational risk management. - Partner with engineering leadership on platform architecture, deployment patterns, and infrastructure automation using Terraform and related infrastructure-as-code tooling. - Design and implement next-generation observability capabilities spanning metrics, logging, tracing, alerting strategy, and actionable dashboards. - Lead major incident response for critical production events, improve escalation paths, and drive durable corrective actions through post-incident reviews. - Influence reliability posture across teams through RFC reviews, engineering guardrails, and resilience engineering work such as disaster recovery, failover design, and capacity forecasting. - Drive continuous improvement of CI/CD systems, reduce operational toil through automation and self-service platform capabilities, and mentor senior engineers and technical leads.
Requirements
- 8+ years of experience in SRE, DevOps, platform engineering, or infrastructure engineering with a strong record of leading large-scale cloud initiatives in production. - Deep expertise in AWS and Kubernetes, including hands-on work designing, operating, and evolving large-scale containerized microservice systems. - Strong experience with Terraform or comparable infrastructure-as-code tooling, and with AWS services for distributed systems and container platforms such as EKS, ECS, or GKE. - Proven success defining and implementing SRE practices such as SLIs, SLOs, error budgets, incident management, observability, and production readiness standards. - Strong background in CI/CD and delivery engineering using tools such as Jenkins, CircleCI, and GitHub Actions, with proficiency in at least one programming language like Python or Go. - Experience operating in security-conscious, regulated environments with familiarity in frameworks such as PCI, SOC 2, and NIST, plus strong communication skills for guiding technical decisions from engineers through executives.
Nice to have
- Experience using AI-assisted and agentic engineering tools to improve productivity, automation, operational insight, and developer workflows with sound judgment around quality, security, and reliability. - A systems-thinking mindset that balances strategic direction with hands-on execution in high-scale, high-ownership environments.