Remote job
Senior Software Engineer, GPU Cluster Infrastructure
Job details
About this role
Role overview
Operate and improve a growing GPU computing platform that supports large-scale research workloads across multiple providers. The role spans Kubernetes operations, scheduling, storage, reliability, and security, partnering directly with researchers to solve infrastructure problems and turn recurring issues into lasting platform improvements.
Responsibilities - Manage the day-to-day lifecycle of Kubernetes GPU clusters, including upgrades, driver and image rollouts, safe rollback, and capacity planning. - Operate batch scheduling and multi-tenant resource controls, including queues, quotas, priorities, preemption, gang scheduling, and fair sharing. - Design and maintain shared and object storage for datasets and checkpoints, including quotas and backups. - Improve training reliability through node health checks, automated draining, investigation of GPU and network issues, and checkpoint and restart practices. - Strengthen platform security through access controls, network policies, secrets management, workload isolation, and sandboxing for AI agents. - Bring new provider capacity into service by testing fabric, training communication, and storage performance; integrate clusters with infrastructure as code and contribute to runbooks, incident reviews, and on-call coverage.
Requirements - At least three years in systems or infrastructure engineering on production Linux, supporting GPU, HPC, or large-scale batch platforms; experience owning a system from design through operation. - Production Kubernetes experience for GPU workloads with a batch scheduler such as Slurm, Kueue, or Volcano, including quotas, priorities, preemption, and node health. - Experience managing infrastructure as code and observability for a production fleet, using tools such as Terraform or Ansible, Helm and ArgoCD, and Prometheus or equivalents. - Strong programming skills in a language commonly used for infrastructure, such as Python, Go, Rust, or C++, and experience maintaining shared automation or services. - Clear written communication for design documents, incident summaries, and coordination with engineering teams, researchers, and providers.
Nice to have - Experience with multi-node training, PyTorch and NCCL troubleshooting, GPU drivers, or InfiniBand and RoCE networking. - Distributed filesystems or object storage at scale, security controls and sandboxing, Kubernetes scheduler internals, or cross-provider platforms.
Benefits and work setup - Full-time work with remote options in many countries and in-person options in Berkeley or Singapore; overlap with Berkeley working hours is preferred. The team currently handles incidents during working hours and expects this role to join an on-call rotation as the platform grows. - Eligible full-time US employees receive health insurance with 94% of the premium paid, a 401(k) match up to 2%, 25 days of annual PTO, up to 10 sick days, paid leave, and a work-from-home stipend and equipment. Office meals are provided in Berkeley. International benefits vary.