Remote job
Research Scientist - Human-AI Systems
Job details
About this role
Role overview A frontier AI research team is hiring a Research Scientist to push forward how high-quality datasets and agentic environments are built. The role sits at the intersection of applied ML research, data engineering, and production systems, with a focus on turning human expertise into scalable, reusable pipelines that close performance gaps for advanced models. Work spans designing experiments, partnering with domain experts, and shipping research into production workflows.
Responsibilities - Design, implement, and optimize reusable pipelines that combine AI capabilities with expert judgment to accelerate data and agentic environment creation. - Run rigorous experiments, including ablations, to validate proof-of-concept approaches and measure impact on data quality, pipeline efficiency, and model performance. - Partner with engineering, data operations, domain experts, and customers to convert research prototypes into reliable production workflows. - Collaborate directly with subject-matter experts to design and test workflows for authoring, reviewing, and refining data and environments. - Work cross-functionally with product, engineering, and data operations to surface findings that inform roadmap decisions. - Stay current with research on synthetic data, agent environments, and data-centric AI, and bring best practices back into internal workflows. - Represent the team's research externally through publications, blog posts, conference talks, and customer engagements.
Requirements - Strong research background in AI, machine learning, NLP, LLMs, or related fields, with experience developing and evaluating new methods. - Experience building environments for AI agents in areas such as automated research, computer use, coding, or professional domain workflows. - Hands-on work with one or more of synthetic data generation, human-in-the-loop workflows, reinforcement learning, agent environments, or model evaluation. - Strong experimental design skills, including hypothesis definition and ablation studies. - Solid software engineering practices, including clean code, modular design, and version control. - Ability to collaborate with domain experts and translate their knowledge into concrete tasks, evaluation criteria, and repeatable workflows. - Comfort with rapid iteration, ambiguous research questions, and moving ideas from experiment to production. - Ph.D. in machine learning, NLP, or a related field preferred; equivalent industry or research lab experience considered.
Benefits and work setup Remote-friendly with options to be based in San Francisco or New York City. Compensation is determined by skills, qualifications, experience, and location, with a published salary range of $200,000–$375,000 USD.