Remote job
Member of Technical Staff, Post-Training
Job details
About this role
Role overview Help shape a new technical organization focused on post-training for frontier AI models. As a Member of Technical Staff on the Post-Training team, you will design and build evaluation frameworks, training environments, data pipelines, and reusable platforms that turn research insight into reliable systems. The role spans the full post-training loop, from hypothesis formation through experiment design and production-grade implementation, working alongside AI researchers, domain experts, and external partners.
Responsibilities - Design post-training systems and methodologies for frontier models, including supervised fine-tuning, reinforcement learning, preference optimization, and reward modeling - Translate open-ended research or partner needs into clear hypotheses, experiments, evaluation plans, and production-quality implementations - Build and improve evaluation frameworks, benchmarks, training environments, data-processing pipelines, and quality-control systems - Run rapid iteration loops: prototype, evaluate, interpret results, and feed learnings back into systems and products - Partner directly with AI researchers and domain experts to develop high-signal data, feedback, and evaluation methods - Identify repeatable patterns across engagements and productize them into reusable software and platforms - Raise the technical bar through clear documentation, design review, and collaboration
Requirements - 3+ years of demonstrated strength in post-training, fine-tuning, or model-evaluation work, including RL, SFT, LoRA/PEFT, full fine-tuning, RLHF, DPO, PPO, or reward modeling - Strong Python skills with the ability to write clean, efficient, and scalable software - Hands-on experience with PyTorch and large-scale data, training, or evaluation workflows - Sound experimental judgment: form hypotheses, choose meaningful metrics, diagnose failures, and separate signal from noise - Experience designing systems, not just implementing specifications, with tradeoffs around quality, scale, reliability, and reuse - Comfort operating in an ambiguous, fast-moving environment with substantial ownership - Collaborative, low-ego communication style with researchers, engineers, domain experts, and customers
Nice to have - Building or operating large-scale ML training, inference, data, or evaluation systems - Developing LLM or agent benchmarks, evaluation methodologies, annotation systems, or data-quality frameworks - Research or applied work on reinforcement learning, alignment, model behavior, synthetic data, or human-in-the-loop systems - Published research, meaningful open-source contributions, or evidence of technical leadership in ML or AI - Experience productizing research or repeated customer work into reusable platforms
Benefits and work setup - San Francisco and Mountain View preferred; open to exceptional candidates in other locations - Equity in a fast-growing company - 401(k) match, competitive compensation, and financial coaching - Paid parental leave, fertility benefits, and parental coaching - Medical, dental, and vision coverage plus mental health support and a wellness stipend - Learning stipend for ongoing development - Commuting support, free lunch, and on-site gym in the San Francisco office - Flexible PTO, 15 holidays, and 2 flex days - Team outings and referral bonuses