Remote job
Staff+ Site Reliability Engineer, Safeguards ML Infra
Job details
About this role
Role overview
Senior-plus site reliability engineer responsible for the reliability, scalability, and operability of machine learning infrastructure supporting safety safeguards. The role blends classic SRE disciplines with the unique demands of large-scale production ML systems. Periodic travel is expected to collaborate with on-site engineering hubs.
Responsibilities
- Own SLOs and incident response for ML infrastructure supporting safeguards workloads - Design and improve observability, capacity planning, and deployment pipelines for large model serving systems - Partner with ML engineers and researchers to make training and inference platforms reliable and cost-efficient - Lead post-incident reviews and drive systemic remediations across teams - Contribute to on-call rotations and operational tooling for ML infrastructure
Requirements
- Extensive experience operating large-scale distributed systems in an SRE or production engineering capacity - Familiarity with ML infrastructure concerns such as GPU scheduling, model serving, and data pipelines - Strong scripting or software engineering skills in languages commonly used for infrastructure tooling - Track record of leading complex cross-team reliability initiatives - Willingness to travel periodically to major engineering hubs
Nice to have
Background supporting safety, moderation, or policy-related ML systems.
Benefits and work setup
Remote-friendly with periodic travel to San Francisco, Seattle, or New York offices.