Remote job
TypeScript Engineer, AI Coding Agent Evaluator
Job details
About this role
Role overview This contract role focuses on evaluating how contemporary AI coding agents perform in realistic developer situations. The engineer serves as a senior reviewer, judging whether the model's responses reflect sound engineering judgment, useful reasoning, and the kind of trustworthy guidance an experienced developer would expect.
Responsibilities - Evaluate AI-generated coding interactions end to end and judge whether each response is useful, broadly correct, and aligned with how a strong engineer would think. - Assess the quality of explanations, preambles, and reasoning alongside the produced code. - Distinguish between levels of response quality and articulate what makes an answer weak versus strong. - Provide clear, opinionated feedback describing what worked, what did not, and what felt misleading or off. - Help shape quality standards for AI-assisted coding experiences, including tools such as Cursor.
Requirements - Staff or Principal-level engineering background or equivalent real-world experience. - Strong TypeScript or JavaScript expertise, with the ability to evaluate code without executing or line-by-line reviewing it. - Hands-on experience with AI coding tools such as OpenAI Codex, Claude Code, or Cursor. - Deep familiarity with AI-assisted developer workflows and how engineers use them in practice. - Comfort making subjective but rigorous judgments and communicating them directly. - High bar for what good engineering looks like in both code and conversation.
Nice to have - Experience with AI-first IDEs and related developer tooling. - Prior exposure to prompt design, evaluation rubrics, or model review processes. - Background mentoring senior engineers or setting engineering standards across a team.
Benefits and work setup - Contract engagement at $100 to $200 per hour. - Around 10 to 20 hours per week, starting ASAP and running through early May, with possible extension. - Hiring process consists of a take-home evaluation exercise followed by a single behavioral interview.