Remote job
SWE-Bench AI Task Auditor - Freelance AI Trainer Project
Job details
About this role
Role overview
A freelance, remote contract position for software engineering specialists who audit tasks used to train and evaluate AI systems. The role focuses on validating that tasks are technically sound, realistic, solvable, and reproducible within the SWE-Bench specialty.
Responsibilities
- Review software engineering tasks for technical accuracy, realism, solvability, and reproducibility. - Evaluate accompanying tests and evaluation criteria for reliability. - Identify codebase integration issues, test failures, and logic errors within task designs. - Provide clear, actionable feedback that helps task authors improve quality. - Apply real-world application development knowledge to judge whether tasks reflect genuine engineering scenarios.
Requirements
- Demonstrable professional experience in software engineering. - Strong ability to navigate complex codebases. - Familiarity with SWE-Bench standards and real-world application development. - Analytical and problem-solving skills for testing and troubleshooting complex technical scenarios.
Benefits and work setup
Pay rate of $60 per hour, with the final rate determined after evaluating experience and geographic location. Engagement is a freelance contract, fully remote. Contractors supply their own secure computer and high-speed internet; company-sponsored health insurance and PTO do not apply.