Remote job
Senior SRE
Job details
About this role
Role overview
The Senior SRE owns operational excellence for a portfolio of modernized, distributed SaaS applications running across multiple cloud environments. Working as part of a team that provides round-the-clock coverage, the role blends Tier 1 reliability engineering, observability, disaster recovery, and security incident response on AWS and Azure. It is a hands-on, code-intensive position that leans on Infrastructure-as-Code and AI-driven automation to reduce toil at scale.
Responsibilities
- Participate in rotating on-call coverage for containerized production applications across multiple business units, ensuring continuous availability and rapid response - Maintain and extend observability tooling (monitoring, logging, tracing) to detect degradation early and shorten mean-time-to-detect and mean-time-to-resolve - Develop, test, and execute disaster recovery and business continuity procedures for cyber, infrastructure, or geographic disruptions - Respond to security incidents against SaaS platforms by following established runbooks and coordinating remediation - Build and operate Infrastructure-as-Code (Terraform) and CI/CD pipelines (GitHub Actions, GitLab CI) to automate deployments and reduce manual work - Design and operate AI agents that automate SRE tasks such as triage, detection, and remediation within a DevSecOps practice - Act as a technical escalation point, applying strong analytical skills to resolve infrastructure, network, and automation issues across distributed multi-tenant SaaS environments
Requirements
- 5–7 years of progressive experience in software engineering and/or site reliability engineering, focused on operating distributed systems - Deep coding expertise in Python, JavaScript, or Go, including APIs, authentication, parallelization, triggering, and data transformation - Strong hands-on experience with container technologies (Docker, Kubernetes) supporting highly scalable and resilient distributed systems - Production Terraform experience working with modules at scale - Operational experience with AWS (for example EC2, Lambda, EKS, S3, RDS) and/or Azure (such as Container Apps, AKS, Container Storage) - Deep CI/CD background with GitHub Actions or GitLab CI and a track record of embedding DevSecOps practices into operational workflows - Experience operating highly available production systems with application-level logging, troubleshooting, and tracing tools
Nice to have
- Hands-on experience building or operating AI agents that automate SRE tasks and incident response - Familiarity with security incident response procedures and runbook-driven operations
Benefits and work setup
- Remote role open to candidates based in the United States or Canada - US salary band of approximately USD 130,000–165,000 and Canada band of CAD 120,000–145,000, excluding annual bonus and equity where applicable - Compensation varies based on market conditions, location, job-related skills, and experience