Remote job
Engineering Manager, SRE
Job details
About this role
Role overview
Guide a Site Reliability Engineering team whose mission is to let product engineers move quickly while customers experience a consistently available service. This position blends roughly 60% hands-on technical work with 40% people leadership, reporting to the Director of Engineering for Platform and managing four direct reports. You'll set direction, develop engineers, and represent the team across the broader engineering organization.
Responsibilities
- Run the full people lifecycle for direct reports, covering onboarding, feedback, performance assessment, progression, and hiring - Shape team goals and prioritization, and design the support rotation and on-call model - Own core platform infrastructure: Kubernetes, AWS, PostgreSQL, DNS and TLS, and CI - Build out the reliability practice, including SLOs, error budgets, incident response, and the observability stack - Work closely with security on threats, patching, infrastructure controls, and audit and compliance requirements - Manage vendor relationships and renewals, with the Director supporting commercial conversations
Requirements
- Demonstrated leadership of an SRE, infrastructure, or platform engineering team, with ownership of career growth and progression - Deep hands-on background in site reliability, DevOps, or cloud infrastructure engineering - Production Kubernetes experience, including the messy operational reality beyond the happy path - AWS at meaningful scale, infrastructure as code with Terraform, and Docker plus shell scripting - CI/CD systems such as GitLab CI, GitHub Actions, or Jenkins - Solid observability practices and a history of running incident response, on-call, SLOs, and error budgets, ideally in regulated environments - Clear written communication suited to an asynchronous, distributed team
Nice to have
- Backend language experience, ideally Elixir, or otherwise Java, Clojure, Node.js, or Python - Observability depth with OpenTelemetry, distributed tracing, or tools such as Honeycomb - PostgreSQL or Aurora operations: performance tuning, connection pool health, query analysis - Hands-on experience building or scaling AI infrastructure - Security capability across defensive and offensive domains, plus FinOps and cloud cost management
Benefits and work setup
- Fully remote with work-from-anywhere flexibility and asynchronous-first working hours - Flexible paid time off - 16 weeks paid parental leave - Budget for co-working, learning, and wellness, including gym memberships - Mental health support services, stock options, and a home office plus IT equipment budget - Annual base salary range of $75,450 to $169,700 USD, dependent on location, transferable skills, and experience