Remote job
Post-Training & RLHF Engineer
Job details
About this role
Role overview This position is part of an AI Research team dedicated to advancing post-training methods for modern machine learning models, with a focus on Reinforcement Learning from Human Feedback (RLHF). The engineer will help design, run, and refine the pipelines that turn pre-trained models into aligned, capable systems. The role is full-time and fully remote.
Responsibilities - Build and iterate on RLHF and related post-training pipelines for large models - Develop reward models, preference datasets, and evaluation harnesses - Partner with researchers to translate experimental ideas into reproducible training runs - Monitor training dynamics, diagnose regressions, and surface alignment issues - Maintain clean code, documentation, and experiment tracking practices
Requirements - Practical experience with RLHF or comparable post-training techniques - Strong Python skills and familiarity with modern ML training frameworks - Understanding of reward modeling, preference data, and alignment evaluation - Comfort designing experiments, analyzing metrics, and debugging training runs
Benefits and work setup - Full-time, remote position