Remote job
Technical Lead - GPU Infrastructure (100% Remote - Worldwide)
Job details
About this role
Role overview
Lead the architecture and delivery of a GPU compute and managed inference platform that runs on Kubernetes, taking the system from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure. This is a 100% remote, globally distributed role that combines hands-on technical leadership with line management of about twelve engineers across backend, frontend, DevOps, QA, and documentation.
Responsibilities
- Own platform architecture end to end: high-level and low-level designs, review-driven proposals, and a kept-current baseline. - Lead and line-manage a distributed engineering team, setting standards, running code and design review, release gates, one-to-ones, and growth input. - Design, build, and operate a managed Slurm scheduling layer for research users, including controller and accounting, partitions and login nodes, driver and CUDA baselines, node onboarding, stalled-job and node-health detection, drain, and autohealing. - Own the Kubernetes control plane and GPU enablement on partner-provided bare metal: NVIDIA GPU Operator and Network Operator, KubeVirt/VFIO-based VM GPU isolation, and day-2 operations including upgrades, backup and recovery, and node replacement. - Architect managed inference at scale: multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute capacity for sensitive workloads. - Set up observability and operations across control plane, GPU fleet, and application tiers, with SLOs, incident response, post-incident review, and an on-call model a small team can sustain. - Act as the primary technical interface to infrastructure partners and vendors, translating requirements into written specifications and acceptance tests, running escalations to closure, and contributing to capacity planning and hardware sourcing. - Translate internal research, model-training, and product team workloads into platform requirements and broker capacity when it is constrained. - Complete the platform team hiring and set the technical bar for the engineers who join it.
Requirements
- 8+ years of hands-on engineering, including at least three years leading teams that build and operate infrastructure platforms other teams depend on; Bachelor's or Master's in computer science or engineering, or equivalent practical experience. - Hands-on Slurm at scale: operational experience running slurmctld and slurmdbd for real users, including partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system; ideally operated an HPC or GPU training cluster for a research population. - GPU fleet operation on bare metal: NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, and node burn-in and acceptance. - High-performance interconnects: InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems. - Linux systems depth: kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, and performance tuning for compute-heavy workloads. - Production Kubernetes operation (not just deployment): control plane, upgrades, CNI and CSI, operators and custom controllers, and multi-tenancy design. - HPC storage and data movement experience plus infrastructure-as-code and GitOps.
Nice to have
- Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab's platform team. - Peer-to-peer or distributed-systems background. - Experience working with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.