Remote job
Senior Site Reliability Engineer
Job details
About this role
Role overview Join a Platform SRE team to build and operate the infrastructure, tools, and paved roads that help developers deliver scalable, secure, and reliable software. The role spans infrastructure automation, observability, developer enablement, and system reliability, with a focus on reducing toil and creating self-service capabilities across the engineering organization.
Responsibilities - Design, build, and scale production environments using AWS and Terraform, driving architectural decisions that improve long-term maintainability - Lead efforts to improve platform resilience through failure-based testing, automated recovery strategies, and proactive capacity planning - Own the design and delivery of reusable platform components and self-service tools that streamline the developer experience - Define and evolve observability standards including system metrics, distributed tracing, and SLO frameworks - Drive projects end to end from scoping and estimation through planning, execution, and rollout - Mentor engineers, uphold infrastructure quality, and shape best practices and standards used across the organization - Engage in technical design discussions, provide guidance, and adapt strategies based on team input - Participate in a low-volume on-call rotation
Requirements - 6+ years of experience in SRE, DevOps, Cloud Engineering, or Software Development roles - Hands-on experience operating production environments in AWS - Proficiency in Go or Python with experience building production-grade automation, tooling, or libraries - Strong experience with Terraform for infrastructure as code - Experience with container orchestration platforms such as ECS or Kubernetes - Familiarity with CI/CD tools such as GitHub Actions - Solid understanding of observability practices including system metrics, distributed tracing, and SLOs - Experience with failure-based testing approaches and automated recovery strategies - Strong written and verbal communication and leadership skills
Nice to have - Experience with microservices architectures - Exposure to Kafka or other event streaming systems - Background building internal developer platforms or self-service infrastructure - Familiarity with systems security, compliance requirements, or hardening practices
Benefits and work setup - 100% medical, dental, and vision coverage - Flexible paid time off - Annual home office stipend and WeWork access - Mental and physical health wellness programs - Remote-first, highly inclusive culture with opportunities for advancement - Competitive compensation, with U.S. base salary for this position ranging from $151,000 to $201,000 depending on location, and adjusted ranges for other Canadian markets