Remote job
Senior Research Scientist, STEM
Job details
About this role
Role overview This position focuses on STEM-focused research that advances how frontier AI systems are evaluated, trained, and improved. The work targets problems where the right benchmark, dataset, or methodology does not yet exist, and where researchers must propose new directions from scratch. Projects span hypothesis generation, experimentation, benchmark construction, and publication across frontier STEM evaluation, synthetic data, hallucination and reliability, and agentic science.
Responsibilities - Identify important gaps in benchmark and evaluation literature and design novel benchmarks within and across STEM domains. - Develop rigorous task-generation, grading, contamination-control, difficulty-calibration, and validation methodologies that set new standards for evaluating frontier models. - Create methods for generating high-quality synthetic STEM training data and run experiments that separate genuine capability gains from raw training volume. - Investigate hallucination, uncertainty, calibration, and epistemic failure, and build evaluations for factual reliability, self-correction, verification, citation, and appropriate abstention. - Research long-horizon scientific agents and design workflows combining literature search, coding, simulation, tool use, experimentation, and iterative reasoning. - Pursue new research directions in reasoning, model evaluation, AI-for-science, data generation, and emerging capabilities.
Requirements - PhD or equivalent research experience in a highly technical field such as machine learning, computer science, mathematics, physics, chemistry, biology, engineering, or statistics. - Demonstrated ability to formulate and execute original research projects. - Strong grounding in modern large language models and the frontier AI research landscape. - Excellent experimental design and quantitative analysis skills.
Benefits and work setup - Compensation range of $250,000 to $350,000 OTE plus equity. - Opportunity to publish at leading conferences such as ICLR, ICML, and NeurIPS. - High autonomy, rapid iteration, and meaningful impact on frontier AI research.