Remote job
Research Scientist / Engineer – Training Infrastructure
Job details
About this role
Role overview
This position sits at the intersection of research and infrastructure engineering, focused on building the distributed training systems that power large-scale multimodal foundation models running across thousands of GPUs. The work spans hard PyTorch and CUDA engineering alongside distributed-systems design, with an emphasis on advanced parallelism, training stability, and high utilization across massive clusters. It is built for engineers who have already solved real foundation-model training problems rather than those still ramping into distributed systems.
Responsibilities
- Design, implement, and optimize distributed training systems that scale to thousands of GPUs while preserving throughput and stability - Research and integrate advanced parallelization strategies, including FSDP, tensor, pipeline, and expert parallelism - Build monitoring, visualization, and debugging tooling that makes large training runs observable and diagnosable end to end - Tune training stability, convergence behavior, and resource utilization across massive GPU fleets
First 90 days
- Days 1–30 — Immerse and diagnose the current training stack, identifying where stability and utilization break down at scale - Days 30–60 — Ship a parallelization or stability improvement that measurably helps a real production training run - Days 60–90 — Build the monitoring and reliability tooling that keeps multi-thousand-GPU runs healthy and efficient
Requirements
- Extensive hands-on experience with distributed PyTorch training and the parallelisms used in foundation-model development - Deep understanding of GPU cluster architecture, including networking and storage subsystems - Familiarity with communication libraries such as NCCL and MPI, paired with a track record of distributed-system optimization
Nice to have
- Strong Linux systems administration and scripting skills - Experience managing training runs across 100+ GPU deployments - Background with containerization, orchestration, and cloud infrastructure