Remote job
Engineering Manager, SRE
Job details
About this role
Role overview
Lead a Site Reliability Engineering team that keeps a global HR and employment platform running reliably for customers worldwide. The role is roughly 60% individual contributor and 40% people leadership, reporting to the Director of Engineering for Platform with four direct reports. You'll shape reliability practices that are still maturing while staying close enough to the technical work to set direction with credibility and recognize trouble early.
Responsibilities
- Own the full career lifecycle of direct reports, including onboarding, feedback, performance assessment, progression, and hiring - Set SRE goals and prioritization, and manage the support rotation and on-call model - Steward core infrastructure: Kubernetes, AWS, PostgreSQL, DNS and TLS, and CI systems - Drive the reliability practice forward, including SLO rollout, error budgets, incident response, and the observability stack - Partner with the security team on threats, patching, infrastructure controls, and audit and compliance obligations - Manage vendor relationships and renewals behind the platform with support from the Director
Requirements
- Prior leadership of an SRE, infrastructure, or platform engineering team, with real ownership of reports' growth and progression - Hands-on background in site reliability, DevOps, or cloud infrastructure, deep enough to review work and challenge designs - Production Kubernetes experience, including the operational reality rather than the happy path - AWS at meaningful scale, plus infrastructure as code with Terraform - CI/CD experience with systems such as GitLab CI, GitHub Actions, or Jenkins - Docker, shell scripting, and a track record running incident response, on-call, SLOs, and error budgets in a regulated environment
Nice to have
- Working knowledge of a backend language (Elixir preferred, otherwise Java, Clojure, Node.js, Python, or similar) - Modern observability depth with OpenTelemetry, distributed tracing, or tools such as Honeycomb - PostgreSQL or Aurora performance, connection pool health, and query tuning - Security capability from both defensive and offensive standpoints - Cloud cost management and FinOps, plus experience growing a team from a small base
Benefits and work setup
- Fully remote with work-from-anywhere flexibility and an asynchronous-first culture - Flexible paid time off and flexible working hours - 16 weeks paid parental leave and mental health support services - Stock options, a learning budget, and a home office plus IT equipment budget - Budget for local co-working spaces or in-person social events - Annual base salary range of $75,450 to $169,700 USD, dependent on location, skills, and experience