Remote job
Senior Infrastructure Operations Engineer (EMEA)
Job details
About this role
Role overview A senior infrastructure operations role focused on keeping large-scale GPU and bare-metal infrastructure reliable, performant, and ready for production AI workloads. The position spans the full infrastructure lifecycle, from provisioning through decommissioning, with a strong emphasis on automation, root-cause analysis, and cross-team collaboration. It serves as a senior escalation anchor when complex, cross-stack incidents arise in a production environment operating at meaningful scale.
Responsibilities - Operate and troubleshoot large GPU and bare-metal fleets across Linux, compute, networking, storage, and cluster orchestration layers. - Act as a senior escalation point for complex infrastructure incidents, driving investigation through to resolution and root cause analysis. - Own end-to-end operational workflows including provisioning, configuration, validation, maintenance, remediation, and decommissioning. - Build automation and internal tooling that eliminate repetitive operational work and let the fleet scale efficiently. - Identify recurring failure modes and partner with engineering teams to design more reliable, repeatable, and automated systems. - Improve provisioning, monitoring, diagnostics, and operational processes as the infrastructure footprint grows. - Participate in a distributed primary/secondary on-call rotation supporting production infrastructure.
Requirements - Significant experience operating and troubleshooting Linux-based production infrastructure at scale. - Strong automation and scripting skills using Python, Go, Bash, Ansible, or comparable tools. - Hands-on experience with Kubernetes, Slurm, or other cluster and workload orchestration platforms. - Working knowledge of monitoring, observability, and telemetry systems for diagnosing production issues. - Solid systems and networking fundamentals, with a proven track record of owning production problems through to resolution. - Comfort working in ambiguous environments and collaborating across engineering, operations, and customer-facing teams.
Nice to have - Experience troubleshooting, provisioning, or operating bare-metal or high-speed data center networking. - Familiarity with hardware management and provisioning technologies such as PXE, BMC, IPMI, Redfish, or iDRAC. - Background with distributed or high-performance storage systems such as VAST, Ceph, GPFS, or WEKA. - Experience with infrastructure-as-code, configuration management, or GitOps workflows.
Benefits and work setup - Comprehensive medical, dental, and vision coverage for employees and eligible dependents. - Equity participation, plus a discretionary bonus component for eligible roles. - Retirement savings support, including 401(k) matching in the U.S. and pension contributions in the U.K. - Unlimited PTO, company holidays, and floating holidays, plus a two-week company-wide winter break. - Paid parental and family leave, plus a four-week paid sabbatical after four years of service. - Annual learning and development allowance, wellness stipends, and work-from-home support. - Flexible schedules with a hybrid work model for office-based teams and complimentary in-office meals at office hubs.