← Back to jobs

Remote job

Technical Lead - GPU Infrastructure (100% Remote - Worldwide)

DevOps Full-time Worldwide

Job details

Not specified Salary
Worldwide Eligibility
Lead Experience
Full-time Employment

About this role

Role overview

Lead the architecture and delivery of a GPU compute and managed inference platform that runs on Kubernetes, taking the system from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure. This is a 100% remote, globally distributed role that combines hands-on technical leadership with line management of about twelve engineers across backend, frontend, DevOps, QA, and documentation.

Responsibilities

- Own platform architecture end to end: high-level and low-level designs, review-driven proposals, and a kept-current baseline. - Lead and line-manage a distributed engineering team, setting standards, running code and design review, release gates, one-to-ones, and growth input. - Design, build, and operate a managed Slurm scheduling layer for research users, including controller and accounting, partitions and login nodes, driver and CUDA baselines, node onboarding, stalled-job and node-health detection, drain, and autohealing. - Own the Kubernetes control plane and GPU enablement on partner-provided bare metal: NVIDIA GPU Operator and Network Operator, KubeVirt/VFIO-based VM GPU isolation, and day-2 operations including upgrades, backup and recovery, and node replacement. - Architect managed inference at scale: multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute capacity for sensitive workloads. - Set up observability and operations across control plane, GPU fleet, and application tiers, with SLOs, incident response, post-incident review, and an on-call model a small team can sustain. - Act as the primary technical interface to infrastructure partners and vendors, translating requirements into written specifications and acceptance tests, running escalations to closure, and contributing to capacity planning and hardware sourcing. - Translate internal research, model-training, and product team workloads into platform requirements and broker capacity when it is constrained. - Complete the platform team hiring and set the technical bar for the engineers who join it.

Requirements

- 8+ years of hands-on engineering, including at least three years leading teams that build and operate infrastructure platforms other teams depend on; Bachelor's or Master's in computer science or engineering, or equivalent practical experience. - Hands-on Slurm at scale: operational experience running slurmctld and slurmdbd for real users, including partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system; ideally operated an HPC or GPU training cluster for a research population. - GPU fleet operation on bare metal: NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, and node burn-in and acceptance. - High-performance interconnects: InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems. - Linux systems depth: kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, and performance tuning for compute-heavy workloads. - Production Kubernetes operation (not just deployment): control plane, upgrades, CNI and CSI, operators and custom controllers, and multi-tenancy design. - HPC storage and data movement experience plus infrastructure-as-code and GitOps.

Nice to have

- Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab's platform team. - Peer-to-peer or distributed-systems background. - Experience working with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Skills detected in the listing

JavaScriptReactNode.jsKubernetesLLM
Detected Oct 9, 2026
Last verified Oct 9, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight