Remote job
Senior GPU Systems Engineer
Job details
About this role
Role overview This is a senior systems engineering position focused on running large-scale GPU fleets that power customer AI training and inference workloads. The role spans the full lifecycle of GPU hardware, from commissioning new racks to diagnosing performance regressions under multi-node load, working closely with platform and ML infrastructure teams.
Responsibilities - Commission new GPU racks end to end, including burn-in, hardware validation, firmware and driver baselines, topology verification, and fleet acceptance. - Diagnose failures at the hardware boundary, covering GPU faults, NVLink and NVSwitch degradation, PCIe and RDMA fabric issues, and thermal and power events. - Tune node-level performance for real workloads through NUMA placement, CPU and interrupt affinity, huge pages, and GPUDirect and RDMA paths. - Build and maintain automated health-check and burn-in tooling so that bad nodes are quarantined before customers schedule onto them. - Own driver, CUDA toolkit, and firmware version policy across the fleet, including staged rollout and rollback procedures. - Investigate multi-node training and inference performance regressions alongside the ML infrastructure team and drive them to root cause. - Carry the on-call rotation for fleet health and author runbooks for novel failure modes.
Requirements - Substantial production experience operating large-scale GPU or accelerated-compute infrastructure. - Deep familiarity with NVIDIA H100 hardware, with exposure to B200 or next-generation silicon considered relevant. - Strong systems programming skills in Python, Go, or Bash, with a discipline of treating tooling as a first-class deliverable. - Hands-on experience carrying production on-call for infrastructure with real availability commitments. - Clear written communication skills, since a significant portion of the role involves documenting findings so the next on-call engineer does not relearn them.
Nice to have - Experience bringing up a brand-new GPU generation, including issues that only appear on new silicon. - Familiarity with Kubernetes device plugins, GPU operators, or scheduling GPUs in containerised environments. - Background in data-centre-side concerns such as rack power budgeting, liquid cooling, or high-density thermal design. - Contributions to relevant open-source projects or published benchmarking work.