Remote job
Member of Technical Staff, ML Platform
Job details
About this role
Role overview
This role sits within an ML platform team at a frontier AI research organization building world simulation models. The mission is to own the end-to-end evaluation platform that determines whether models are ready to ship. The work spans tooling, automation, and standards across research teams working on video, image, audio, agents, and robotics.
Responsibilities
- Design and own the evaluation platform, including tooling for sample generation, automated scoring, and human annotation at scale - Define CLIs, APIs, GUIs, and storage layers that make evaluation seamless, fast, and reproducible - Embed directly with research teams to understand domain-specific measurement needs and generalize solutions - Establish standards for reproducibility, metric definitions, reporting formats, and trustworthy results - Drive platform adoption across training, production model serving, and other research efforts - Contribute broadly to ML platform infrastructure that supports frontier model development
Requirements
- 5+ years building production ML infrastructure or data platforms, with some experience in evaluation, experimentation, or benchmarking systems - Strong Python and PyTorch, plus hands-on experience running large batch GPU workloads on Kubernetes - Experience designing data pipelines and storage for large volumes of media or model outputs, with attention to versioning and reproducibility - Familiarity with experimental statistics, including paired comparisons, confidence intervals, multiple-comparison pitfalls, and inter-rater agreement - Comfort building internal tools end to end, from the command line to the browser - Ability to lead a broad technical area by gathering requirements, setting direction, and driving a roadmap independently
Nice to have
- Track record building evaluation suites for generative models - Hands-on experience with LLM- or VLM-as-judge pipelines - Familiarity with online experimentation platforms and connecting offline metrics to product outcomes - Prior work evaluating agents or robotics policies