Remote job
AI Runtime + Evals Engineer
Job details
About this role
Role overview Own the runtime and evaluation systems that keep production AI agents dependable. You’ll build the infrastructure for agent execution and recovery, then measure whether changes improve task success, cost, and latency across client-specific workflows.
Responsibilities - Build and maintain agent runtime capabilities for tool calls, context handling, streaming, and recovery when a run fails. - Improve resilience across model providers with fallbacks, retries, cancellation, and cost controls. - Develop evaluation datasets, automated graders, CI regression gates, and evaluations over production traces. - Track task success, cost, and latency, including separate evaluation sets for different clients. - Investigate intermittent and provider-specific failures, malformed tool schemas, and context overflow. - Make context handling token-aware through techniques such as offloading, summarization, and budget limits.
Requirements - Experience operating agentic systems in production, including tool use, context management, streaming, and failure recovery. - Experience evaluating agent systems with datasets, graders, metrics, and CI gates. - Strong debugging skills for probabilistic, intermittent, and model-provider-specific issues. - Strong general software engineering ability; familiarity with TypeScript is useful, and experience in other stacks can transfer. - Strong written and spoken English.
Nice to have - Experience with LangGraph, LangChain, or similar agent frameworks, or low-level integrations with a major model provider. - Awareness of inference costs and security considerations.
Benefits and work setup - Full-time, remote role with direct engineering ownership. - Work on production agent systems in a demanding, fast-moving environment.