Job details
About this role
Role overview
This is a senior individual contributor role focused on the reliability, scalability, and modernization of a cloud platform serving large enterprise customers with demanding availability expectations. The position blends deep hands-on infrastructure work with technical leadership across a distributed systems engineering function, helping shape architecture across multiple cloud providers and modern container-orchestrated environments.
Responsibilities
- Drive infrastructure modernization, evolving legacy compute environments toward container-orchestrated infrastructure while keeping enterprise service stable - Design and maintain infrastructure as code across multiple cloud providers with portability and long-term scalability in mind - Improve CI/CD systems and deployment tooling so releases are efficient, observable, and recoverable - Provide technical leadership through mentorship, architecture guidance, and knowledge sharing across the team - Establish and evolve SLI and SLO practices, alongside monitoring, alerting, and load-testing capabilities - Lead during platform incidents, including troubleshooting, root cause analysis, and follow-through on corrective actions - Drive cloud cost visibility and embed cost considerations into architecture decisions - Improve developer experience by evolving infrastructure, deployment workflows, and production feedback loops - Partner with Security and Engineering on access controls, hardening, and compliance requirements
Requirements
- 8+ years of hands-on experience in infrastructure, DevOps, platform engineering, or site reliability engineering - Deep experience designing, operating, and troubleshooting highly available production infrastructure - Production experience across more than one major cloud provider, with depth in at least one - Hands-on experience building, operating, or significantly improving CI/CD systems and deployment infrastructure - Strong incident response skills for complex distributed-system failures, including effective post-incident review - Strong Linux administration skills plus scripting or programming ability in Ruby, Python, or a comparable language - SaaS background where reliability, availability, and production stability are critical - Ability to collaborate across US and European time zones and participate in an on-call rotation - Strong communication skills with the ability to mentor engineers and influence infrastructure decisions - Bachelor's degree in a related field or equivalent practical experience
Nice to have
- Operating production infrastructure across both AWS and Google Cloud simultaneously - Prior experience leading or managing engineers, including in a player-coach capacity - Cloud cost management or FinOps practices at meaningful scale - Managing deployment platforms such as Spinnaker, Jenkins, or comparable tooling - Operating a Ruby on Rails enterprise application or comparable production codebase - Familiarity with SOC 2 or similar compliance frameworks and customer-facing security requirements - Working knowledge of AI-assisted development tooling and its infrastructure implications - Background in learning management systems or adult education platforms
Benefits and work setup
- 100% employee premium coverage for selected medical, dental, and vision plans - 401(k) with matching for US-based employees - Flexible paid time off - LinkedIn Learning access - Calm subscription - Annual company-wide retreat - Remote-first organization with team members distributed globally - Equal opportunity employer committed to building an inclusive team