Remote job
Research Engineer, Model Evaluations
Job details
About this role
Role overview
This role builds the evaluations and supporting infrastructure used to measure what advanced language models can do. You will turn broad questions about reasoning, knowledge, agent behavior, personality, and safety into defensible metrics, then run those evaluations reliably against models during training. The work combines experimental design, dataset and scoring development, distributed systems, observability, and clear communication of results.
Responsibilities
- Design evaluations for model capabilities and safety properties, including tasks that require reasoning or agentic behavior. - Create visualizations and reporting that make evaluation results understandable to researchers and decision-makers. - Build and harden distributed execution systems so large evaluation suites run reliably against training checkpoints. - Own dashboards used to monitor model health, improving signal quality, latency, and regression visibility. - Investigate anomalous results during training and distinguish model changes from problems in data, harnesses, or infrastructure. - Improve libraries, workflows, and tooling so research teams can implement and iterate on evaluations efficiently.
Requirements
- Strong software-engineering ability and experience building reliable systems or research infrastructure. - Experience with machine learning, evaluation methodology, data processing, or large-scale experimentation. - Ability to define measurable success criteria for ambiguous capabilities and validate that metrics reflect meaningful behavior. - Clear communication under time pressure, especially when diagnosing unexpected results. - Flexibility across team boundaries and willingness to take ownership of unassigned but important work. - Comfort with collaborative development, including pair programming.
Nice to have
Experience with large-scale dataset sourcing, curation, and processing; ML training infrastructure; distributed evaluation systems; or experiments involving prompting, sampling, and scaffolding is useful.
Benefits and work setup
The role is remote-friendly but requires travel and may be based from offices in San Francisco or New York City. The general expectation is that staff spend at least 25% of their time in an office, with some roles requiring more. Visa sponsorship may be available depending on the role and candidate. The listed annual salary range is $500,000–$850,000 USD.