Remote job
Senior Site Reliability Engineer / DevOps
Job details
About this role
Role overview Drive the reliability and operational excellence of production platforms that power cloud services and AI workloads. This senior role combines Site Reliability Engineering and DevOps to keep systems fast, resilient, and observable through rigorous SLO work, deep automation, and disciplined incident response.
Responsibilities - Define, instrument, and enforce SLOs and SLIs across cloud and AI workloads, using error budgets to guide engineering trade-offs. - Build and maintain automation for deployments, scaling, and routine operations to reduce toil and improve reproducibility. - Lead incident response end-to-end, including detection, triage, root-cause analysis, blameless post-mortems, and follow-through on corrective actions. - Advance observability through thoughtful logging, metrics, and tracing, ensuring teams can answer unknown questions about production behavior. - Partner with product and platform engineers to design services that meet resilience targets from day one.
Requirements - Several years of hands-on experience as an SRE or DevOps engineer operating production cloud environments at scale. - Strong working knowledge of Kubernetes, infrastructure-as-code tooling, and CI/CD pipelines. - Practical experience with monitoring and observability stacks (metrics, logs, traces) and incident management practices. - Familiarity with at least one major cloud provider and Linux systems engineering. - Solid scripting or programming skills for automation and tooling.
Nice to have - Exposure to AI/ML inference platforms, GPU-backed workloads, or model-serving infrastructure. - Experience with chaos engineering, capacity planning, or cost optimization.