Remote job
Site Reliability Engineer
Job details
About this role
Role overview Build and operate the reliability, delivery, automation, and security foundations that allow product teams to release software quickly and safely. Working within a productivity engineering and self-service platform team, this role owns production reliability practices, incident response, observability, infrastructure automation, Kubernetes operations, deployment safety, and security controls embedded in CI/CD.
Responsibilities - Define and evolve SLI, SLO, and error-budget practices, using reliability data to influence priorities and product decisions. - Lead incident response, facilitate post-incident reviews, and convert findings into durable system improvements. - Build and maintain observability across metrics, logs, and traces while improving signal quality and reducing alert fatigue. - Design and operate resilient infrastructure with Infrastructure as Code, including capacity planning and cloud-cost optimization. - Manage production Kubernetes and container workloads and support safe deployment methods such as canary releases, progressive rollouts, and rapid rollback. - Integrate and tune SAST, DAST, SCA, dependency scanning, and related security controls in delivery pipelines. - Implement policy-as-code to prevent unsafe infrastructure and Kubernetes changes at admission time. - Maintain vulnerability triage and remediation service levels, improve on-call sustainability, and coach engineers on operational practices.
Requirements - 5+ years in site reliability, platform, or infrastructure engineering with senior ownership of production systems. - Strong programming ability in Go, Python, TypeScript, or a similar language for automation and production tooling. - Hands-on experience with a major cloud platform, Kubernetes, and Infrastructure as Code; AWS and Terraform experience are useful. - Proven experience leading incident response and implementing SLO-driven reliability practices. - Working knowledge of observability tooling; experience with Datadog is useful. - Practical experience securing CI/CD through scanning, dependency controls, or policy-as-code. - Understanding of cloud security fundamentals, including IAM, least privilege, guardrails, and secrets management. - Strong judgment and communication skills when raising reliability or security issues across engineering teams.
Nice to have - Experience with policy-as-code frameworks such as OPA/Rego, Kyverno, or Conftest.
Benefits and work setup The source describes comprehensive healthcare, retirement savings with employer matching, paid family leave, fertility support, mental health resources, wellness and technology allowances, performance-related rewards, learning opportunities, mentorship, and an inclusive workplace culture.