Remote job
Python Engineer, AI Coding Agent Evaluator
Job details
About this role
Role overview This contract engagement centers on evaluating the behavior of modern AI coding agents in realistic developer scenarios. The engineer acts as a senior reviewer, judging whether model outputs reflect sound engineering judgment, useful reasoning, and the kind of guidance an experienced developer would trust.
Responsibilities - Review AI-generated coding interactions end to end and assess whether each response is useful, broadly correct, and aligned with how a strong engineer would think. - Judge the quality of explanations, preambles, and reasoning, not just the code produced. - Calibrate between levels of response quality and explain what separates a mediocre answer from a strong one. - Deliver clear, opinionated feedback on what worked, what missed, and what felt misleading or off. - Contribute to quality standards for AI-assisted coding experiences, including tools like Cursor.
Requirements - Staff or Principal-level engineering experience, or equivalent depth from real-world work. - Strong Python background, with the ability to evaluate code without needing to execute or deeply review every line. - Hands-on experience with AI coding tools such as OpenAI Codex, Claude Code, or Cursor. - Solid grasp of AI-assisted developer workflows and the ways engineers actually use them. - Comfort making rigorous subjective judgments and expressing them directly. - High standards for what good engineering looks like in both code and written reasoning.
Nice to have - Prior use of AI-first IDEs and similar developer tooling. - Familiarity with prompt design, evaluation rubrics, or model review workflows. - Experience mentoring senior engineers or defining engineering standards across teams.
Benefits and work setup - Contract role paying $100 to $200 per hour. - Approximately 10 to 20 hours per week, starting ASAP and running through early May, with potential extension. - Process includes a take-home evaluation exercise and a single behavioral interview.