← Back to jobs

Remote job

HPC Infrastructure Engineer - GPU Clusters

DevOps Full-time Permanent US

Job details

Not specified Salary
US Eligibility
Staff Experience
Full-time Employment

About this role

Role overview This role supports large-scale AI research by owning the GPU cluster infrastructure that powers model training. The engineer operates at the intersection of systems engineering and HPC, building and running production Linux and GPU environments end-to-end. It is a hands-on, high-impact position on a small team that values automation, root-cause fixes, and direct collaboration with researchers.

Responsibilities - Maintain and evolve the stack beneath training code, including OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, and high-speed networks such as InfiniBand or RoCE - Operate and tune job schedulers like Slurm so researchers get fair, fast access to compute - Build and care for high-performance storage for datasets and checkpoints - Diagnose and remediate performance issues across nodes, degraded links, thermal behavior, and flaky GPUs, fixing the class of problem rather than just the instance - Run preventive health checks, automated draining, remediation, and burn-in pipelines for new capacity - Benchmark rented GPU capacity, validate SLAs, and step in for hands-on hardware work such as racking, cabling, and diagnostics when needed

Requirements - Experience running large-scale Linux server or GPU environments in production, with comfort in both building and operating them - Strong familiarity with the NVIDIA stack (drivers, CUDA, NCCL, DCGM) or transferable deep systems experience - Comfort with bare-metal hardware, server components, and high-speed networking - Solid automation skills in Python and/or Bash, plus infrastructure-as-code tools such as Ansible or Terraform - Ability to dig through metrics, logs, and PromQL to identify the real cause of issues - Willingness to own scope end-to-end and pitch in on any task, including datacenter trips

Nice to have - Experience supporting distributed ML training from the infrastructure side, including failure modes and checkpointing patterns - Background evaluating GPU cloud providers and using parallel filesystems like WEKA or VAST, or large-scale object storage - BMC/IPMI/Redfish automation, PXE provisioning at scale, and power and cooling awareness for dense GPU deployments

Benefits and work setup - Fully remote role that can be executed globally, with the option to work from offices in London, New York, San Francisco, and Warsaw - Small, high-trust team with minimal bureaucracy and short feedback loops between infrastructure and researchers - Aggressive automation philosophy aimed at keeping on-call sustainable and preventing repeat pages - Annual discretionary stipend for learning and development, plus a separate annual stipend for team social travel - Monthly co-working stipend for those outside major hubs - Annual company offsite bringing the full team together in a new location each year

Skills detected in the listing

PythonTerraform
Detected Sep 3, 2026
Last verified Sep 3, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight