Remote job
Senior Infrastructure Software Engineer
Job details
About this role
Role overview Join an infrastructure engineering team responsible for the software systems that operate and manage large-scale GPU and bare-metal compute. This software-first role combines backend engineering with Linux systems knowledge to build automation that bridges physical infrastructure with the platforms customers rely on for AI/ML training, inference, and HPC workloads. The work focuses on making fleet-wide infrastructure management more reliable, repeatable, and programmable across thousands of servers, storage systems, and high-performance networks.
Responsibilities - Design, build, and operate production services, APIs, tooling, and automation that manage bare-metal and GPU infrastructure at scale - Develop systems supporting the full infrastructure lifecycle, including provisioning, configuration, monitoring, maintenance, and decommissioning - Build server discovery, provisioning, configuration, validation, and lifecycle management systems that automate fleet bring-up and capacity deployment - Integrate infrastructure software with networking, storage, and data center operations to bring new capacity online efficiently - Translate recurring operational issues and hardware failure modes into durable improvements in software and tooling - Partner with Network, Infrastructure Operations, Data Center, and Platform Engineering teams to define requirements and shape technical direction, including design documents and engineering best practices
Requirements - 8+ years of professional software engineering, infrastructure engineering, or related experience - Strong software engineering fundamentals with production backend systems built in Python or a comparable object-oriented language - Solid Linux experience in production environments - Hands-on experience building APIs, tooling, or automation for managing infrastructure at scale - Familiarity with containerization and orchestration concepts - Understanding of HPC and bare-metal infrastructure fundamentals, including provisioning and out-of-band management - Comfort navigating ambiguity and making pragmatic architecture decisions in a fast-paced, high-ownership environment
Nice to have - Bare-metal hardware troubleshooting and provisioning experience with technologies such as PXE/iPXE, BMC, Redfish, or IPMI, particularly with Dell hardware - Experience with GPU servers in bare-metal or virtualized environments - Experience with network switches, routers, and firewalls, particularly SONiC switches, Palo Alto firewalls, or Juniper Networks - Familiarity with high-performance storage systems, particularly VAST - Prior experience supporting AI/ML or HPC infrastructure at scale
Benefits and work setup - Anticipated annual base salary range of $180,000–$220,000 USD, plus discretionary bonus and equity - Comprehensive medical, dental, and vision coverage for employees and eligible dependents - 401(k) matching (U.S.) and pension contributions (U.K.) - Unlimited PTO, company holidays, floating holidays, and a two-week company-wide winter break - Paid parental and family leave - Annual professional development allowance - Wellness and work-from-home stipends - Four weeks of paid sabbatical after four years of service - Flexible schedules with a hybrid model for office-based teams and complimentary meals at office hubs - Fully remote within the U.S. or hybrid at designated office hubs; occasional team and company offsites - Visa sponsorship is not available for this role