Remote job
Site Reliability Engineer
Job details
About this role
Role overview
Build and improve the reliability, observability, and operational practices of production platforms. The work spans prevention and response: setting reliability targets, finding risks before they cause outages, and helping teams recover and learn when incidents occur.
Responsibilities - Define and track service-level indicators and objectives for critical services. - Build monitoring and observability with metrics, tracing, logs, and alerting. - Lead incident reviews and turn findings into changes that improve recovery time. - Create runbooks, escalation paths, and production operating procedures. - Assess capacity and failure risks through planning, load tests, and failure analysis. - Automate recurring operational work and partner with engineering teams on production readiness.
Requirements - Significant production experience in SRE, platform engineering, or infrastructure reliability. - Hands-on expertise with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry. - Experience improving distributed-system reliability, including load balancing, failover, circuit breakers, and retries. - Strong incident response, on-call, root cause analysis, and post-incident review practices. - Ability to automate operational tasks with Python, Bash, Go, or similar scripting tools. - Clear communication of reliability needs and tradeoffs to engineering and product partners.
Nice to have - Experience with chaos engineering or failure injection. - SLO or SLA management in enterprise or regulated environments. - Multi-region or multi-cloud reliability experience.
Benefits and work setup - Full-time role, based in Toronto or remote within Canada.