← Back to jobs

Remote job

Senior GPU Systems Engineer

DevOps Remote

Job details

Not specified Salary
Remote Eligibility
Senior Experience
Not specified Employment

About this role

Role overview This is a senior systems engineering position focused on running large-scale GPU fleets that power customer AI training and inference workloads. The role spans the full lifecycle of GPU hardware, from commissioning new racks to diagnosing performance regressions under multi-node load, working closely with platform and ML infrastructure teams.

Responsibilities - Commission new GPU racks end to end, including burn-in, hardware validation, firmware and driver baselines, topology verification, and fleet acceptance. - Diagnose failures at the hardware boundary, covering GPU faults, NVLink and NVSwitch degradation, PCIe and RDMA fabric issues, and thermal and power events. - Tune node-level performance for real workloads through NUMA placement, CPU and interrupt affinity, huge pages, and GPUDirect and RDMA paths. - Build and maintain automated health-check and burn-in tooling so that bad nodes are quarantined before customers schedule onto them. - Own driver, CUDA toolkit, and firmware version policy across the fleet, including staged rollout and rollback procedures. - Investigate multi-node training and inference performance regressions alongside the ML infrastructure team and drive them to root cause. - Carry the on-call rotation for fleet health and author runbooks for novel failure modes.

Requirements - Substantial production experience operating large-scale GPU or accelerated-compute infrastructure. - Deep familiarity with NVIDIA H100 hardware, with exposure to B200 or next-generation silicon considered relevant. - Strong systems programming skills in Python, Go, or Bash, with a discipline of treating tooling as a first-class deliverable. - Hands-on experience carrying production on-call for infrastructure with real availability commitments. - Clear written communication skills, since a significant portion of the role involves documenting findings so the next on-call engineer does not relearn them.

Nice to have - Experience bringing up a brand-new GPU generation, including issues that only appear on new silicon. - Familiarity with Kubernetes device plugins, GPU operators, or scheduling GPUs in containerised environments. - Background in data-centre-side concerns such as rack power budgeting, liquid cooling, or high-density thermal design. - Contributions to relevant open-source projects or published benchmarking work.

Skills detected in the listing

PythonGoAWSKubernetes
Detected Sep 28, 2026
Last verified Sep 28, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight