Remote job
Model Evaluation Engineer — Benchmarks & Evals
AI Engineer Full-time Remote
Job details
Not specified Salary
Remote Eligibility
Not specified Experience
Full-time Employment
About this role
Role overview
Remote, full-time Model Evaluation Engineer role on an AI Research team. The position centers on designing, running, and analyzing benchmarks and evaluations that measure machine learning model behavior and progress.
Responsibilities
- Design and maintain benchmark suites and evaluation methodologies - Run experiments, analyze results, and communicate findings to the AI Research team - Iterate on evaluation infrastructure to support ongoing model development
Requirements
- Experience with ML model evaluation, benchmarking, or research engineering - Strong analytical skills and rigor in experimental design - Comfort operating as part of a fully remote, distributed team