Remote job
Infrastructure Engineer (GPU & Compute)
Job details
About this role
Role overview This position focuses on bringing up, validating, and operating large-scale bare-metal compute environments with an emphasis on GPU-enabled systems. The engineer will sit at the intersection of hardware, systems software, and automation, owning diagnostics, qualification, and the tooling that keeps clusters ready for demanding AI/ML and HPC workloads.
Responsibilities - Own and evolve image management, deployment, and validation pipelines across bare-metal infrastructure, including firmware, driver, and OS qualification for GPU-enabled systems - Operate and maintain test clusters used for bring-up and diagnostics, supporting hardware qualification efforts for next-generation platforms - Diagnose and resolve complex issues spanning GPUs, drivers, OS, and underlying hardware; analyze performance using tools such as NVIDIA DCGM - Build Python-based automation for provisioning, validation, and system bring-up, improving reliability, repeatability, and scalability - Manage Linux-based production and validation environments, including virtualization and PXE/image-based bare-metal provisioning workflows - Partner with infrastructure, hardware, and data center teams, plus platform and ML stakeholders, to ensure systems meet workload requirements and contribute to provisioning and lifecycle best practices
Requirements - 5+ years in infrastructure engineering, systems engineering, or a closely related role - Strong Linux systems experience in production environments - Hands-on experience with GPU-enabled systems and diagnostic tooling such as NVIDIA DCGM - Familiarity with bare-metal provisioning and system bring-up workflows - Proficiency in Python or comparable scripting/programming languages for automation - Ability to debug complex issues across hardware, OS, GPUs, and system software
Nice to have - Experience with high-performance interconnects such as InfiniBand or NVLink - Familiarity with PXE boot environments, LiveCD systems, or image-based provisioning workflows - Working knowledge of hardware management interfaces (iDRAC, IPMI, Redfish) - Data center operations experience with physical hardware - Background supporting AI/ML or HPC workloads at scale - Experience with GPU validation frameworks or large-scale hardware qualification processes
Benefits and work setup - Anticipated annual base salary range of $180,000–$220,000 USD, plus discretionary bonus and equity - Comprehensive medical, dental, and vision coverage for employees and eligible dependents - Retirement savings support (401(k) matching in the U.S., pension contributions in the U.K.) - Unlimited PTO, company holidays, floating holidays, and a two-week company-wide winter break - Paid parental and family leave, plus a four-week paid sabbatical after four years of service - Annual professional development allowance, plus wellness and work-from-home stipends - Flexible schedules with a hybrid model for office-based teams and complimentary in-office meals at hubs - Work may be fully remote within the U.S. or hybrid out of office hubs in NYC, SF, Seattle, or London, with occasional team and company offsites; visa sponsorship is not available for this role