Remote job
Distributed Training Researcher
Job details
About this role
Role overview
This is a full-time research position on a small in-house lab that publishes work on multi-thousand-GPU pre-training, long-context modeling, and inference acceleration. The lab sits between the runtime engineering team and customer-facing teams, with a mandate to make the underlying platform faster, more affordable, and capable of longer contexts in writing. Researchers lead one to two papers per year with full compute backing and at least one engineer co-author, and each publication is expected to produce a follow-up change in production.
Responsibilities
- Lead one to two papers per calendar year as first or co-first author, choosing the research question with full compute backing - Run experiments at the 256-to-2048 GPU scale on B200 and B300 hardware, including long-context regimes up to one million tokens - Translate findings into runtime pull requests alongside the inference-acceleration team, with every paper expected to yield a production-impact follow-up - Mentor one to two early-career researchers each year, including PhD students from partner universities - Represent the lab at major venues such as NeurIPS, ICML, ICLR, MLSys, and ASPLOS, and contribute to at least one open benchmark per quarter - Publish independently under a publication veto, with no organizational delay or editing of papers
Requirements
- PhD in computer science, electrical and computer engineering, or applied mathematics, or an equivalent publication record with three or more first-author papers at top ML or systems venues - Hands-on experience designing and running experiments at thousand-GPU scale - Strong PyTorch skills plus production-scale comfort with at least one of FSDP, DeepSpeed, or Megatron-LM - Empirical, write-up-first research style, including willingness to publish negative results when useful - Willingness to ship code that runs directly on customer workloads rather than working in a pure sandbox
Nice to have
- Existing co-appointment or close ties with a research university - Hands-on experience with long-context architectures such as ring or paged attention, or low-level kernel work in Triton, CUDA, or ROCm
Benefits and work setup
- Remote position open to candidates in the US or EU - Reports to the lab director with a light on-call presence limited to one learning-only shadow rotation per year - Access to B200 and B300 clusters connected by an InfiniBand fabric, with tools including PyTorch, FSDP, DeepSpeed, Megatron-LM, wandb, and aim