Remote job
Senior Site Reliability Engineer / DevOps
Job details
About this role
Role overview A senior platform engineering role focused on keeping production cloud environments fast, resilient, and observable, including the AI workloads running on top of them. The engineer partners closely with client teams to define meaningful SLOs, reduce toil through automation, and lead incident response when things go wrong. It is a hands-on position that shapes how modern, AI-driven services stay dependable at scale.
Responsibilities - Define and operate SLOs, SLIs, and error budgets for production services running on cloud infrastructure alongside AI workloads. - Build and maintain automation for deployment, scaling, and routine operational tasks to remove manual toil. - Lead incident response, including on-call coordination, blameless post-incident reviews, and follow-up reliability work. - Improve observability through thoughtful logging, metrics, and tracing instrumentation across services. - Partner with engineering teams on architecture decisions that affect availability, scalability, and recovery.
Requirements - Significant senior-level experience in Site Reliability Engineering, DevOps, or platform engineering. - Deep familiarity with at least one major cloud provider such as AWS, GCP, or Azure. - Hands-on experience with Infrastructure-as-Code tooling like Terraform, Pulumi, or CloudFormation. - Strong background in incident management, SLO definition, and operational excellence practices. - Proficiency with container orchestration, ideally Kubernetes, in production.
Nice to have - Familiarity with GPU-accelerated compute, inference platforms, or MLOps tooling. - Experience with chaos engineering or resilience testing. - Contributions to open-source observability or reliability projects.