Remote job
Senior Product Engineer (Agent Training and Evals)
Job details
About this role
Role overview Join an early-stage engineering team building infrastructure that captures and evaluates AI agent traces at scale, alongside real human expert work, in secure environments. This is a genuine 0→1 role with significant ownership over technical direction — turning ambiguous problems about training, benchmarking, and evaluating AI agents into scalable systems.
Responsibilities - Build tooling for capturing and processing data from agents and humans performing real-world tasks at significant scale - Solve hard problems around compute, orchestration, scaling, security, and reliability - Help develop approaches for training, benchmarking, and evaluating AI agents - Work with emerging agentic AI technologies and rapidly evolving models - Make pragmatic technical decisions in a highly ambiguous environment - Partner closely with product engineering and AI research to turn customer problems into working products - Drive direct impact within a fast-moving, high-ownership, cross-functional team
Requirements - Strong experience building scalable infrastructure or distributed systems in the cloud - Solid understanding of how systems are deployed, scaled, and secured - Hands-on experience integrating and evaluating AI models or agents - Genuine interest in how AI agents are trained, evaluated, and improved - Strong product mindset with a focus on delivering business and customer value - Comfort with ambiguity and ability to move quickly from idea to working system - Strong communication and collaboration skills
Nice to have - Experience with image, video, or audio processing - Python experience (useful but not essential)
Benefits and work setup - Remote working within a mission-driven culture - Key technologies include Google Cloud Platform, Python, JavaScript, and TypeScript; FastAPI and Vue.js; container-based and serverless architectures; Postgres SQL; GitHub Actions, Kubernetes, and Datadog for DevOps and monitoring