Remote job
Lead Engineer, AI Quality
Job details
About this role
Role overview A fully remote, globally distributed engineering organization is hiring a hands-on Lead Engineer to own the evaluation and observability foundation for its production LLM-powered agent systems. The position combines player-coach technical leadership of an AI Quality engineering team with deep individual contribution: designing the pipelines, scorers, and diagnostic infrastructure that determine whether AI agents are working correctly, improving them, and reducing their cost and latency. This is a build-and-iterate role in a fast-moving environment where shipping experiments and letting data settle arguments is the norm.
Responsibilities - Design and own evaluation infrastructure, including CI/CD pipelines, scorers, and datasets that assess AI agents across single tool calls and full multi-turn conversations. - Trace quality breakdowns across the agent pipeline (planning, tool selection, execution, sub-agents) and translate findings into prioritized fixes or prototypes. - Build and scale annotation workflows, including human-labeled and AI-generated or simulated conversation datasets that expand product coverage faster than manual labeling alone. - Run structured experiments comparing prompts, models, and agent harness configurations against production baselines, then drive decisions on model swaps, caching, and routing strategies to cut cost and latency. - Set technical direction and day-to-day priorities for the AI Quality engineering team while remaining hands-on in the codebase. - Partner closely with the core AI engineering team to give product changes real confidence in quality improvements, not just successful deployment.
Requirements - 7+ years building and shipping production software, including significant work on LLM-powered agents that take real actions inside a product. - Deep experience with complex, tool-using systems involving multiple tools, planning or orchestration, and sub-agents, with the ability to walk through a shipped example and the metrics that proved it worked. - Working proficiency in Ruby on Rails (the production foundation) and Python, or the ability to ramp quickly on both. - Hands-on experience building evaluation or observability infrastructure for ML/AI systems, such as eval pipelines, scorers, dashboards, or eval CI/CD, ideally with established evaluation frameworks. - Experience designing datasets, annotation workflows, or labeling pipelines, including human annotation, AI-generated data, and simulation approaches. - Comfort working in a highly dynamic, fast-paced environment with ambiguity, running many experiments, and embracing the ones that do not work as learning.
Nice to have - Proficiency in English at CEFR Level C2 or ILR Level 5 for spoken, written, and reading contexts.
Benefits and work setup - Fully remote, async-first team with high autonomy and minimal hours monitoring. - Twice-yearly company-wide retreats in international destinations. - U.S.-benchmarked compensation offered globally, plus equity with ongoing refresh grants. - 35 days of paid time off per year. - Benefits package supporting health, wellbeing, and professional growth.