Remote job
Site Reliability Engineer
Job details
About this role
Role overview
This is a senior Site Reliability Engineer (SRE III) role on a team that runs and continuously improves a large, global OpenStack-based hosting platform. The work blends deep production operations with engineering: automating toil, planning capacity, instrumenting systems, and protecting customer experience through migrations and incident response. You will serve as a technical lead others lean on for complex reliability challenges across a multi-region footprint.
Responsibilities
- Operate and scale cloud infrastructure spanning compute, networking, and storage in a large, multi-region production environment. - Lead the safe migration of customer hosting workloads onto the platform, designing migration tooling, validation steps, and rollback strategies. - Eliminate manual operational work through Python-based automation and configuration management, treating repeated effort as a defect. - Strengthen observability by improving monitoring, alerting, and dashboards so signal reaches the right engineer quickly. - Participate in a shared on-call rotation, lead incident response when on call, and run blameless post-incident reviews that turn failures into permanent fixes. - Mentor junior engineers, review peers' designs and code, and maintain clear runbooks and system documentation. - Contribute to internal AI-assisted operational workflows that help engineers resolve problems and incidents faster.
Requirements
- 5+ years in SRE, infrastructure, platform, or systems engineering roles operating production systems at scale. - Strong Linux fundamentals across networking, storage, processes, and performance troubleshooting. - Proficiency in Python for maintainable, tested automation and tooling. - Hands-on experience operating distributed systems and diagnosing issues across service boundaries. - Experience with infrastructure-as-code and configuration management such as Puppet or Ansible. - Demonstrated ownership of production reliability, including on-call participation, incident response, and toil reduction. - Clear written and verbal communication for documentation, post-incident reviews, and remote collaboration.
Nice to have
- Experience operating OpenStack components such as Nova, Neutron, or Ceph, or comparable cloud infrastructure platforms. - Familiarity with Docker, Kolla, or containerised infrastructure. - Experience supporting large distributed systems or operating software-defined storage such as Ceph at scale. - Exposure to OpenStack or broader cloud migration programs. - Interest in applying AI or LLM tooling to operational workflows.
Benefits and work setup
- Fully remote position with occasional in-person team meetings or events. - Compensation package may include paid time off, retirement savings plans, bonus or incentive eligibility, equity grants, and an employee stock purchase plan. - Competitive health benefits plus family-friendly offerings such as parental leave, with specifics varying by location and discussed during the interview process.