Remote job
Software Engineers: Paid Code Review for AI Agent Evaluation
Job details
About this role
Role overview Participate in a paid remote research study aimed at improving how artificial-intelligence coding agents are benchmarked against realistic software-engineering scenarios. You will review programming tasks and their associated evaluation harnesses, judging whether they accurately reflect real-world engineering problems. Your technical feedback will be gathered through a guided conversation and screen-sharing session.
Responsibilities - Review realistic programming tasks for accuracy, difficulty, and relevance to real engineering work. - Verify the logic, test cases, and overall structure of the provided evaluation harnesses. - Assess whether coding environments effectively measure software-engineering skills. - Walk through your thought process out loud while analyzing complex code structures and test setups.
Requirements - Active professional experience working as a software engineer. - Hands-on experience building or maintaining test suites, evaluation harnesses, or comprehensive code reviews. - Comfortable reviewing code and explaining technical concepts verbally during a live session. - Familiarity with complex, real-world software architecture across backend, full-stack, or machine-learning systems.
Nice to have - Background as a backend engineer, full-stack developer, machine-learning engineer, or software architect. - Experience designing or critiquing benchmarks or assessment environments for engineering candidates or AI systems.
Benefits and work setup - Fully remote, single-session engagement with flexible scheduling. - Compensation of $65 per hour for the duration of the study. - Opportunity to influence how AI coding agents are evaluated across the industry.