Remote job
RL AI Research Scientist
Job details
About this role
Role overview An RL AI Research Scientist role on a research team building the reinforcement learning foundation of an AI agent platform. The work targets real-world enterprise tasks, with emphasis on long-horizon planning, reward shaping, and context selection that allow agents to operate reliably at scale. The position is remote, with a preference for candidates based in the US or Singapore.
Responsibilities
- Design and implement novel reinforcement learning algorithms for training agents on complex, multi-step enterprise workflows. - Develop reward modeling, context selection, and policy optimization techniques that improve agent accuracy over extended horizons. - Run large-scale experiments, analyze results rigorously, and translate research into production-ready components. - Collaborate with infrastructure engineers to ensure research prototypes scale efficiently on cloud and on-device hardware. - Track advances in RL, LLM fine-tuning, and agent architectures, and propose new research directions. - Contribute to intellectual property through publications, patents, and open-source work.
Requirements
- PhD or equivalent research experience in Reinforcement Learning, Machine Learning, or a closely related field. - Strong publication record at top venues such as NeurIPS, ICML, ICLR, or AAAI. - Deep expertise in RL fundamentals, including policy gradient methods, value-based methods, model-based RL, multi-agent RL, or RLHF/RLAIF. - Proficiency in Python and at least one deep learning framework, with a strong preference for PyTorch. - Demonstrated ability to move research from prototype to production.
Nice to have
- Experience training and fine-tuning large language models. - Background in on-device or edge inference optimization, including quantization, distillation, or mixture-of-experts architectures. - Familiarity with enterprise software deployment, compliance, or regulated industries. - Open-source contributions in RL or LLM ecosystems. - Experience with distributed training at scale using frameworks such as FSDP, DeepSpeed, or Megatron.