Remote job
Team Leader, SRE
Job details
About this role
Role overview
This role leads the Site Reliability Engineering function that powers a global product platform, owning Kubernetes, AWS, PostgreSQL, the CI infrastructure, and the observability stack that sits on top of them. The position blends roughly 60 percent hands-on technical work with 40 percent people leadership, reporting to the Director of Engineering, Platform, and managing four direct reports. It suits a leader who wants a maturing reliability practice where the foundations are in place and the harder scaling questions are still open.
Responsibilities
- Own the full people lifecycle for the team: hiring, onboarding, feedback, performance reviews, progression, and growth conversations. - Set and sequence the team's commitments against company priorities, balancing operational load against project delivery. - Run the support rotation and on-call model, and step in with credibility during incidents. - Maintain the core platform layer, including Kubernetes, AWS, PostgreSQL, DNS and TLS, and CI infrastructure, including vendor relationships and renewals. - Mature the reliability practice by extending SLO coverage, error budgets, incident response, and observability to the remaining product teams. - Partner closely with Security on threats, patching, infrastructure controls, and audit and compliance obligations. - Act as the team's spokesperson across engineering and to senior leadership, and handle team dynamics and conflict directly.
Requirements
- Demonstrated leadership of an SRE, infrastructure, or platform engineering team, with genuine ownership of career development rather than just sprint delivery. - Hands-on background in site reliability, DevOps, or cloud infrastructure, with production Kubernetes experience including the non-happy-path scenarios. - Operational experience with AWS at meaningful scale, and Terraform-based infrastructure as code. - Practical use of CI/CD systems such as GitLab CI, GitHub Actions, or Jenkins, plus Docker and shell scripting. - Experience running incident response, on-call, SLOs, and error budgets, and turning incidents into durable changes. - Strong written communication, since most leadership happens asynchronously in a distributed team. - Familiarity with regulated environments and audit or compliance work.
Nice to have
- Working knowledge of a backend language, ideally Elixir, or alternatively Java, Clojure, Node.js, or Python. - Depth in modern observability, including OpenTelemetry, distributed tracing, and tools like Honeycomb. - Database operations experience with PostgreSQL or Aurora, including connection pool health and query tuning. - Hands-on Linux work outside cloud environments, security capability from both offensive and defensive angles, and FinOps or cloud cost management.
Benefits and work setup
- Fully remote and async-first, with flexible working hours and the ability to work from anywhere. - Flexible paid time off and 16 weeks paid parental leave. - Mental health support, stock options, learning budget, and home office plus IT equipment allowance. - Budget for co-working spaces, learning, and wellness, including gym memberships. - Annual base salary range of $75,450 to $169,700 USD, dependent on location and level.