Remote job
ML Research Engineer — Model Pre-Training
Job details
About this role
Role overview An ML Research Engineer focused on model pre-training works on an AI research team to advance foundational model development. The position blends research and engineering, designing and running large-scale pre-training experiments that shape downstream model capabilities. It is a full-time remote role embedded in a research-driven environment.
Responsibilities - Design, implement, and execute pre-training experiments across model architectures, scales, and data configurations - Develop and maintain distributed training infrastructure, including compute orchestration and pipeline tooling - Monitor long-running training jobs, diagnose instability, and iterate on data, optimization, and architectural choices - Build evaluation harnesses and benchmark pipelines to measure the quality of pre-trained checkpoints - Collaborate with researchers and engineers to translate experimental findings into concrete model improvements - Contribute to internal documentation, code reviews, and write-ups of research outcomes
Requirements - Hands-on experience training large-scale neural networks with frameworks such as PyTorch or JAX - Familiarity with distributed training techniques, including data, tensor, or pipeline parallelism across GPU or TPU clusters - Solid grasp of transformer architectures, modern optimization methods, and pre-training objectives - Strong Python skills and the ability to write clean, reproducible research code - Comfort working in an iterative research setting where hypotheses are formed and tested continuously
Nice to have - Publications at ML venues such as NeurIPS, ICML, or ICLR, or contributions to open-source ML projects - Experience with data curation, tokenization, or large-scale data pipeline engineering