Remote job
Customer Cluster Engineer
Job details
About this role
Role overview This role sits within a small customer engineering team that pairs senior engineers with reserved-capacity accounts running large-scale AI workloads. Rather than traditional solutions architecture or support, the position owns end-to-end performance for a handful of clients, profiling and tuning the same training and inference stacks the internal runtime team ships. It is a full-time senior individual contributor position based in San Francisco or remote across the US and EU.
Responsibilities - Act as the named technical owner for three to five reserved-capacity accounts running 256–2048-GPU jobs. - Profile distributed training workloads, examining NCCL collectives, gradient overlap, checkpoint cost, and MFU, and defend specific optimization recommendations. - Tune kernels and configurations alongside the runtime team, with changes shipping through the same review process and release train. - Drive technical pre-sales for expansions and renewals, including pilot design, scope, success criteria, and post-mortems. - Lead incident reviews for any P1 affecting owned accounts and route root-cause work into runtime, cluster, or platform teams. - Share the customer-engineering on-call rotation with runtime and cluster SRE, roughly one week in six.
Requirements - Five or more years of distributed-systems or ML-systems engineering, with production experience on 64+ GPU jobs. - Strong PyTorch, FSDP, and Megatron-LM debugging skills, including the ability to read a wandb run and identify where time is going. - Working knowledge of CUDA and at least one of Triton or CUTLASS, sufficient to read kernel code without writing it daily. - Clear written and verbal communication for quarterly post-mortems and customer-facing presentations. - Comfort discussing technical trade-offs with executive audiences.
Nice to have - Prior runtime, SRE, or research engineering background. - Familiarity with InfiniBand fabrics, NCCL tuning, and tooling such as Linux perf or aim. - Experience contributing kernel or configuration changes upstream through a shared release process.