Remote job
Team Leader, SRE
Job details
About this role
Role overview
This is a hybrid leadership and hands-on engineering role leading a Site Reliability Engineering team inside a global Platform Engineering organization. The mission is to keep the core product fast and available for engineers and end users by owning Kubernetes, AWS, PostgreSQL, the CI infrastructure, and the observability stack. It is roughly a 60 percent individual contributor and 40 percent leadership position, reporting to the Director of Engineering, Platform, with four direct reports.
Responsibilities
- Run the full career lifecycle of four direct reports, including onboarding, feedback, performance assessments, progression planning, and hiring. - Set the SRE team's priorities and commitments in alignment with broader company goals, and decide what work comes first. - Own the support rotation and on-call model, protecting the team's ability to deliver against operational obligations. - Steward the core platform: Kubernetes, AWS, PostgreSQL, DNS and TLS, and the CI infrastructure, including vendor renewals and commercial conversations. - Advance the reliability practice, including SLOs, error budgets, incident response, and post-incident follow-through. - Partner with the Security team on threats, patching, infrastructure controls, and audit and compliance obligations. - Represent the team across engineering and to senior leadership, and resolve team dynamics issues directly rather than routing around them.
Requirements
- Prior leadership of an SRE, infrastructure, or platform engineering team with real ownership of growth, performance, and career progression. - Hands-on background in site reliability, DevOps, or cloud infrastructure engineering, including the operational reality of Kubernetes in production. - Meaningful experience operating AWS at scale. - Practical experience with infrastructure as code using Terraform, plus CI/CD systems such as GitLab CI, GitHub Actions, or Jenkins. - Track record running a reliability practice: incident response, on-call, SLOs, error budgets, and converting incidents into lasting improvements. - Strong written communication skills suited to a fully distributed, async working environment. - Experience working in regulated environments with audit and compliance obligations.
Nice to have
- Working knowledge of a backend language, ideally Elixir, or otherwise Java, Clojure, Node.js, or Python. - Depth in modern observability, including OpenTelemetry, distributed tracing, and tools like Honeycomb. - Database operations experience with PostgreSQL or Aurora, including performance and connection pool tuning. - Security capability from both defensive and offensive angles, and cloud cost or FinOps experience.
Benefits and work setup
- Fully remote, async-first work culture with flexible hours. - Flexible paid time off and 16 weeks paid parental leave. - Mental health support services, stock options, learning budget, and home office plus IT equipment allowance. - Budget for local in-person social events or co-working spaces. - Annual base salary range of $75,450 to $169,700 USD depending on location and level.