Remote job
AI / Agent Engineer
Job details
About this role
Role overview
An engineering role focused on shipping production-grade AI agents that handle real operational workflows such as billing, intake, and back-office operations. The position owns the full agent layer — from tool design and retrieval to evaluation, latency budgeting, and graceful fallback handling — rather than research prototypes or demo notebooks. The expectation is that work lands in front of real users and keeps performing there.
Responsibilities
- Design and implement the agent layer powering production workflows, including tool-use interfaces and retrieval pipelines - Build and maintain evaluation suites such as golden sets and regression tests to measure behavior changes across model and prompt iterations - Define latency budgets and fallback flows that keep multi-step agentic runs reliable under real load - Engineer robustness around retries, idempotency, and partial-failure scenarios so jobs complete correctly even when intermediate calls fail - Ship agent features to real users, monitor their behavior, and iterate based on observed production telemetry
Requirements
- Prior experience deploying agentic systems to live users, not just prototype notebooks - Hands-on work with major model provider APIs and structured tool-use patterns - Strong evaluation discipline, including curating golden datasets and writing regression tests - Fluency in both TypeScript and Python for backend and orchestration code - Comfort designing for retry semantics, idempotent operations, and partial failures
Nice to have
- Public writing or open-source contributions such as an evaluation framework, benchmark, or a published agent incident post-mortem