Remote job
Tech Lead Manager, GPU Cluster Infrastructure
Job details
About this role
Role overview
Lead a small infrastructure team while remaining a hands-on engineer responsible for the technical direction and roadmap of a growing, multi-provider GPU cluster platform. The work spans compute, networking, storage, scheduling, and security, with close collaboration with researchers to keep large experiments performant and dependable.
Responsibilities - Set the platform architecture and roadmap, including how clusters are provisioned, scheduled, monitored, and operated across providers. - Design scheduling and storage systems, and investigate failures across nodes, GPUs, network fabrics, and multi-node jobs. - Hire, coach, and develop senior engineers; establish ownership, priorities, and effective team practices. - Define security controls for a shared research platform, including identity, access, workload isolation, and sandboxing for AI agents. - Establish operating practices for incidents, observability, fault tolerance, and postmortems; participate in incident response and expected future on-call work. - Serve as an escalation point for researchers and infrastructure providers, converting repeated issues into platform improvements.
Requirements - At least five years of systems or infrastructure engineering on production Linux, including GPU, HPC, or large-scale batch platforms, with ownership from design through operation. - Leadership experience as a manager, technical lead, or project lead, including setting direction, scoping work, and giving feedback; direct reports are not required. - Production Kubernetes experience for GPU workloads with a batch scheduler, including quotas, priorities, preemption, and node health. - Experience owning infrastructure as code and observability for a production fleet, using tools such as Terraform or Ansible, Helm and ArgoCD, and Prometheus or equivalents. - Strong programming skills in a systems or infrastructure language such as Python, Go, Rust, or C++, and clear technical writing for varied audiences. - Depth in at least one of scheduling, storage, networking, security, or GPU systems, with enough breadth to review designs across the stack.
Nice to have - Experience with GPU training systems and high-speed fabrics; distributed storage and checkpoint I/O; cluster hardening and sandboxed runtimes; scheduler internals; or multi-provider scheduling and storage. - Experience building a team and establishing on-call, incident, and review practices.
Benefits and work setup - Full-time, 40-hour workweek; remote work is available in many countries, with in-person options in Berkeley or Singapore. Working-hour overlap with Berkeley is preferred, and visa sponsorship is available for in-person employees. - For eligible full-time US employees: health insurance with 94% of the premium paid, a 401(k) match up to 2%, 25 days of annual PTO, up to 10 sick days, paid family and other leave, and a work-from-home stipend and equipment. Office meals are provided in Berkeley. Benefits differ for international hires.