← Back to jobs

Remote job

Research Scientist, Interpretability

AI Engineer Remote considered case-by-case; role based in San Francisco

Job details

Not specified Salary
Remote considered case-by-case; role based in San Francisco Eligibility
Lead Experience
Not specified Employment

About this role

Role overview

Investigate how trained language models implement meaningful algorithms and how their internal mechanisms can be understood well enough to support safer, more trustworthy systems. The work focuses on mechanistic interpretability: developing methods, experiments, and infrastructure to connect model parameters and learned features with observable computations. Researchers and engineers contribute jointly to experimental design, software, analysis, and communication of results.

Responsibilities

- Develop techniques for reverse-engineering algorithms represented in language-model weights. - Design and run experiments ranging from fast toy studies to analyses of large models. - Discover, characterize, and analyze interpretable features and computational circuits. - Build infrastructure for large-scale experiments, data processing, and visualization. - Interpret results carefully, including ambiguous or null findings, and refine methods accordingly. - Explain findings to collaborators and broader technical audiences through clear writing and discussion.

Requirements

- A strong record of scientific research in interpretability or another experimental field. - Familiarity with Python, which is required for the role. - Comfort working in an emerging research area where methods and terminology are still developing. - Ability to treat research and engineering as complementary activities. - Willingness to write code, design experiments, analyze evidence, and communicate conclusions. - Collaborative working style and interest in team-based scientific discovery.

Nice to have

- Experience studying neural networks, language models, learned representations, or model internals. - Experience building experimental tooling or visualizations for machine-learning research. - Interest in collaborating across research groups to apply interpretability findings to model development and safety.

Benefits and work setup

The position is primarily based in a San Francisco office, with exceptional remote arrangements considered case by case. The broader work setup expects staff to spend at least 25% of their time in an office; visa sponsorship may be available depending on the role and candidate.

Skills detected in the listing

Python
Detected Sep 19, 2026
Last verified Sep 19, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight