Remote job
HPC/GPU Systems Engineer
Job details
About this role
Role overview This position sits within a Support and Operations team responsible for the day-to-day health of large-scale GPU fleets powering AI infrastructure. It is a hands-on L2/L3 engineering role covering tickets, alerts, hardware faults, GPU nodes, high-performance networking, Linux, and data centre operations, with a clear pathway toward senior scope through exposure to advanced AI environments.
Responsibilities - Own tickets and tasks end-to-end, escalating early when issues exceed scope and producing clean, evidence-rich handovers. - Perform GPU node triage and hardware troubleshooting, interpreting nvidia-smi or DCGM output and system logs, isolating faults across GPU, NIC, and server hardware, and carrying out physical remediation such as reseats and swap testing. - Run fabric and link diagnostics following established runbooks using tools like mlxlink, capture evidence accurately, and hand off cleanly to senior engineering when needed. - Assist with storage and data-path investigations on high-performance platforms, gathering evidence to support deeper diagnosis. - Maintain accurate records across DCIM, inventory, and asset systems, and contribute simple scripts and tooling improvements that reduce manual work. - Serve as the escalation point for onsite data centre operations staff, coordinating smart-hands tasks within scope.
Requirements - 3-4+ years in infrastructure support or support engineering roles within structured, customer-facing environments. - Working knowledge of GPU infrastructure and hands-on hardware troubleshooting experience. - Strong grasp of Linux fundamentals and the ability to read and interpret system logs. - Clear, structured written and verbal communication, treating documentation quality as a core engineering skill. - Discipline and organisation around ticket hygiene, follow-through, and runbook execution.
Nice to have - Familiarity with Kubernetes concepts such as nodes, pods, services, and logs. - Experience with Ansible, Terraform, CI/CD pipelines such as GitHub Actions, or access tooling such as Teleport or Vault. - Progress toward Linux, networking, Kubernetes, cloud, or security certifications.
Benefits and work setup - Remote-first working model with flexible scheduling. - Compensation includes base plus equity, with reviews every 12 months. - Listed base salary range is $100,000–$140,000 USD, with potential bonus, equity, or commission eligibility. - Medical, dental, vision, paid time off, parental leave, and retirement plan participation may be offered.