Remote job
Staff Research Scientist
Job details
About this role
Role overview This is a research-driven position focused on advancing how frontier AI systems are evaluated, trained, and improved. The work centers on problems where existing benchmarks, datasets, or methodologies fall short, and where researchers must originate new directions. Projects move from hypothesis through experimentation, benchmark construction, and publication across areas like frontier evaluation, synthetic data, hallucination and reliability, and agentic science.
Responsibilities - Identify significant gaps in the existing benchmark and evaluation literature and propose ambitious new research programs to address them. - Design and build novel benchmarks spanning STEM fields and general model functionality, with rigorous task generation, grading, contamination control, difficulty calibration, and validation. - Develop methods for generating high-quality synthetic STEM training data and study how selection, diversity, verification, and filtering affect downstream model performance. - Investigate hallucination, uncertainty, calibration, and epistemic failure, and create evaluations that improve factual reliability, self-correction, verification, citation, and appropriate abstention. - Research AI systems capable of extended scientific and technical work, including long-horizon agents that use literature search, coding, simulation, tool use, and iterative reasoning. - Take projects from initial hypothesis through validated results and contribute findings through publications at leading AI venues.
Requirements - PhD or equivalent research experience in machine learning, computer science, mathematics, physics, chemistry, biology, engineering, statistics, or another highly technical field. - Demonstrated ability to formulate and execute original research, with a track record of substantive contributions. - Strong understanding of modern large language models and the frontier AI research landscape. - Excellent experimental design, quantitative analysis, and technical writing skills.
Benefits and work setup - Compensation range of $250,000 to $400,000 OTE plus equity. - Opportunity to collaborate with colleagues from leading technology companies and publish at top conferences such as ICLR, ICML, and NeurIPS. - High level of autonomy, rapid iteration, and meaningful impact on frontier AI research and enterprise deployment.