Remote job
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Job details
About this role
Role overview
A frontier research engineering role focused on the systems that make reinforcement learning practical at truly large scale. The work spans the full RL loop — training, rollout generation, environment execution, and reward computation — distributed across thousands of GPUs, where throughput, stability, and correctness must all hold together at the same time. This is deep systems work for someone who has already post-trained LLMs with RL, not a modeling-only or research-only seat.
Responsibilities
- Design and scale distributed RL post-training systems that orchestrate trainer, rollout, environment, and reward workloads across thousands of GPUs. - Build high-throughput rollout generation integrating inference engines, weight synchronization, and asynchronous or off-policy schemes. - Develop RL environments for agentic, multi-step tasks, including sandboxed code execution, tool use, computer use, and multimodal interaction, reproducible at millions of episodes. - Construct reward infrastructure covering verifiable and programmatic rewards, reward-model serving, LLM-as-judge pipelines, and defenses against reward hacking. - Build evaluation, monitoring, and debugging tooling that keeps long RL runs stable and diagnosable end to end. - Improve training efficiency and stability, and convert new post-training ideas into production runs alongside research collaborators.
Requirements
- Hands-on experience post-training large language models with RL, covering PPO/GRPO-family methods, RLHF, or RLVR at meaningful scale. - Extensive distributed PyTorch training with parallelism techniques such as FSDP, Tensor Parallel, Pipeline Parallel, or Expert Parallel for foundation models. - Background building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents, including sandboxed execution and multi-turn tool use. - Deep familiarity with RL post-training frameworks such as veRL, OpenRLHF, TRL, or Ray orchestration, plus rollout inference engines like vLLM or SGLang. - Strong understanding of GPU clusters, networking, and communication libraries such as NCCL or MPI under mixed training and inference workloads.
Nice to have
- Running RL training across 100+ GPUs, including asynchronous or disaggregated trainer/rollout architectures. - Containerization and orchestration experience with Kubernetes and Ray for large environment fleets and sandboxed workloads. - Research contributions in RL for LLMs, or open-source contributions to RL training frameworks.