Remote job
Member of Engineering (Evaluations)
Job details
About this role
Role overview
Join a remote-first R&D team working toward artificial general intelligence and ensure that research progress translates into real value for users. In this role you will design and implement the evaluation infrastructure and benchmarks used to measure progress on base models and instruction-following models. The mission is to make sure model improvements show up as meaningful gains on real-world software development skills.
Responsibilities
- Research and implement evaluations and benchmarks for both base models and instruction-following models - Collaborate with applied research and product teams to define metrics that capture progress on real-world software development skills - Build the evaluation tooling and infrastructure that the broader research and engineering organization depends on - Plan, discuss, and communicate clearly with peers across research and engineering - Question existing code quality and evaluation policies, and push for better rigor when warranted
Requirements
- Hands-on experience working with large language models - Strong engineering background with programming skills across multiple languages, including Python - Comfortable working on Linux with strong algorithmic skills - Familiar with the full software development life cycle - Good taste, curiosity, and strong critical thinking
Benefits and work setup
- Fully remote role with flexible hours; team spans Europe and North America, with a focus on EMEA and East Coast candidates - Monthly in-person gatherings of three days plus longer offsites twice a year - 37 days per year of vacation and holidays - Health insurance allowance covering you and dependents - Company-provided equipment plus wellbeing, learning, and home office allowances