Remote job
Manager, Site Reliability Engineering
Job details
About this role
Role overview A cloud-native identity security platform company is hiring a hands-on Manager of Site Reliability Engineering to lead its SRE and DevOps engineers supporting production products. This is a working manager role that requires leading people and leading work, from writing and reviewing automation to commanding major incidents and improving observability practices. The initial scope covers the Platform SRE and DevOps team, with expected expansion into a regulated FedRAMP High environment and the broader commercial platform footprint over time.
Responsibilities - Stay hands-on in the environment by reviewing pull requests, validating pipeline changes, tuning monitors and dashboards, running Datadog queries, and troubleshooting production issues alongside engineers - Own availability and performance of production environments across Azure and AWS, including AKS workloads, ingress and networking, data services, messaging, and CDN or WAF layers - Hire, onboard, coach, and develop full-time SRE engineers while directing and managing contractor resources, scoping work, setting quality expectations, and reviewing deliverables - Run one-on-ones, standups, and planning sessions at times that work for engineers distributed across multiple time zones - Carry the pager as part of the rotation, act as incident commander for Sev1 and Sev2 events, and own the post-incident review process end to end - Operate a public status page, define severity levels, escalation paths, and on-call structure, and ensure RCAs are written to a customer-ready standard with preventative actions assigned to owners
Requirements - Experience operating at scale in a regulated environment such as FedRAMP, with comfort growing scope into adjacent platforms - Hands-on expertise with Azure and AWS production environments, including AKS workloads, networking, data services, messaging, CDN, and WAF - Demonstrated ability to lead a distributed team of full-time engineers and contractors across multiple time zones - Strong incident management experience, including sev definitions, escalation paths, on-call structure, and post-incident review processes - Track record of reducing customer-detected incidents through improved monitoring coverage and meaningful cost optimization across Azure and AWS - Existing US work authorization is required, with no visa sponsorship available now or in the future
Nice to have - Experience with public status page operations and customer notification practices - Familiarity with Atlassian Jira Service Management, Confluence, and Azure DevOps as the operational toolchain
Benefits and work setup - Competitive base salary with a meaningful performance bonus program - Healthcare insurance, pension or retirement matching, comprehensive life insurance, and an employee assistance program - Time off plans and paid company holidays - Inclusive culture built around collaboration, respect, ownership, and global thinking