Remote job
Technical Lead - GPU Infrastructure
Job details
About this role
Role overview
Lead the architecture and delivery of a GPU compute and managed inference platform that runs on Kubernetes, taking the system from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure. The role combines hands-on technical leadership with line management of a distributed team of about twelve engineers across backend, frontend, DevOps, QA, and documentation.
Responsibilities
- Own platform architecture end to end: high-level and low-level designs, review-driven proposals, and a living baseline document. - Lead and line-manage a distributed engineering team, setting engineering standards, running code and design review, release gates, one-to-ones, and growth input. - Design, build, and operate a managed Slurm scheduling layer for research users, covering the controller and accounting, partitions and login nodes, driver and CUDA baselines, node onboarding, stalled-job and node-health detection, drain, and autohealing. - Own the Kubernetes control plane and GPU enablement on partner-provided bare metal, including NVIDIA GPU Operator and Network Operator, KubeVirt/VFIO-based VM GPU isolation, and day-2 operations such as upgrades, backup and recovery, and node replacement. - Architect managed inference at scale: multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity for sensitive workloads. - Set up observability and operations across control plane, GPU fleet, and application tiers, including SLOs, incident response, post-incident review, and an on-call model a small team can sustain. - Act as the primary technical interface to infrastructure partners and vendors, turning requirements into written specifications and acceptance tests, running escalations, and contributing to capacity planning and hardware sourcing. - Translate workloads from internal research, model-training, and product teams into platform requirements and broker capacity when it is short. - Complete the platform team hiring and set the technical bar for the engineers who join it.
Requirements
- 8+ years of hands-on engineering, including at least three years leading teams that build and operate infrastructure platforms other teams depend on; Bachelor's or Master's in computer science or engineering, or equivalent practical experience. - Hands-on Slurm at scale: operational experience running slurmctld and slurmdbd for real users, including partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system; ideally has operated an HPC or GPU training cluster for a research population. - GPU fleet operation on bare metal: NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, and node burn-in and acceptance. - High-performance interconnects: InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems. - Linux systems depth: kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, and performance tuning for compute-heavy workloads. - Production Kubernetes operation (not just deployment): control plane, upgrades, CNI and CSI, operators and custom controllers, and multi-tenancy design. - HPC storage and data movement experience, including shared filesystems, node-local NVMe caching, and distributing large model weights and datasets across many nodes, plus infrastructure-as-code and GitOps.
Nice to have
- Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab's platform team. - Peer-to-peer or distributed-systems background. - Experience working with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.