← Back to jobs

Remote job

Research Engineer, Model Evaluations

AI Engineer Remote-friendly within a US-based role with required travel and possible 25% office attendance

Job details

Not specified Salary
Remote-friendly within a US-based role with required travel and possible 25% office attendance Eligibility
Staff Experience
Not specified Employment

About this role

Role overview

This role builds the evaluations and supporting infrastructure used to measure what advanced language models can do. You will turn broad questions about reasoning, knowledge, agent behavior, personality, and safety into defensible metrics, then run those evaluations reliably against models during training. The work combines experimental design, dataset and scoring development, distributed systems, observability, and clear communication of results.

Responsibilities

- Design evaluations for model capabilities and safety properties, including tasks that require reasoning or agentic behavior. - Create visualizations and reporting that make evaluation results understandable to researchers and decision-makers. - Build and harden distributed execution systems so large evaluation suites run reliably against training checkpoints. - Own dashboards used to monitor model health, improving signal quality, latency, and regression visibility. - Investigate anomalous results during training and distinguish model changes from problems in data, harnesses, or infrastructure. - Improve libraries, workflows, and tooling so research teams can implement and iterate on evaluations efficiently.

Requirements

- Strong software-engineering ability and experience building reliable systems or research infrastructure. - Experience with machine learning, evaluation methodology, data processing, or large-scale experimentation. - Ability to define measurable success criteria for ambiguous capabilities and validate that metrics reflect meaningful behavior. - Clear communication under time pressure, especially when diagnosing unexpected results. - Flexibility across team boundaries and willingness to take ownership of unassigned but important work. - Comfort with collaborative development, including pair programming.

Nice to have

Experience with large-scale dataset sourcing, curation, and processing; ML training infrastructure; distributed evaluation systems; or experiments involving prompting, sampling, and scaffolding is useful.

Benefits and work setup

The role is remote-friendly but requires travel and may be based from offices in San Francisco or New York City. The general expectation is that staff spend at least 25% of their time in an office, with some roles requiring more. Visa sponsorship may be available depending on the role and candidate. The listed annual salary range is $500,000–$850,000 USD.

Skills detected in the listing

PythonStakeholder Management
Detected Sep 19, 2026
Last verified Sep 19, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight