← Back to jobs

Remote job

Tech Lead Manager, GPU Cluster Infrastructure

DevOps Full-time Freelance Worldwide

Job details

$225K – $325K $225K - $325K Salary
Worldwide Eligibility
Lead Experience
Full-time Employment

About this role

Role overview

Lead a small infrastructure team while remaining a hands-on engineer responsible for the technical direction and roadmap of a growing, multi-provider GPU cluster platform. The work spans compute, networking, storage, scheduling, and security, with close collaboration with researchers to keep large experiments performant and dependable.

Responsibilities - Set the platform architecture and roadmap, including how clusters are provisioned, scheduled, monitored, and operated across providers. - Design scheduling and storage systems, and investigate failures across nodes, GPUs, network fabrics, and multi-node jobs. - Hire, coach, and develop senior engineers; establish ownership, priorities, and effective team practices. - Define security controls for a shared research platform, including identity, access, workload isolation, and sandboxing for AI agents. - Establish operating practices for incidents, observability, fault tolerance, and postmortems; participate in incident response and expected future on-call work. - Serve as an escalation point for researchers and infrastructure providers, converting repeated issues into platform improvements.

Requirements - At least five years of systems or infrastructure engineering on production Linux, including GPU, HPC, or large-scale batch platforms, with ownership from design through operation. - Leadership experience as a manager, technical lead, or project lead, including setting direction, scoping work, and giving feedback; direct reports are not required. - Production Kubernetes experience for GPU workloads with a batch scheduler, including quotas, priorities, preemption, and node health. - Experience owning infrastructure as code and observability for a production fleet, using tools such as Terraform or Ansible, Helm and ArgoCD, and Prometheus or equivalents. - Strong programming skills in a systems or infrastructure language such as Python, Go, Rust, or C++, and clear technical writing for varied audiences. - Depth in at least one of scheduling, storage, networking, security, or GPU systems, with enough breadth to review designs across the stack.

Nice to have - Experience with GPU training systems and high-speed fabrics; distributed storage and checkpoint I/O; cluster hardening and sandboxed runtimes; scheduler internals; or multi-provider scheduling and storage. - Experience building a team and establishing on-call, incident, and review practices.

Benefits and work setup - Full-time, 40-hour workweek; remote work is available in many countries, with in-person options in Berkeley or Singapore. Working-hour overlap with Berkeley is preferred, and visa sponsorship is available for in-person employees. - For eligible full-time US employees: health insurance with 94% of the premium paid, a 401(k) match up to 2%, 25 days of annual PTO, up to 10 sick days, paid family and other leave, and a work-from-home stipend and equipment. Office meals are provided in Berkeley. Benefits differ for international hires.

Skills detected in the listing

PythonGoRustC++KubernetesTerraform
Detected Oct 9, 2026
Last verified Oct 9, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight