Remote job
AI Engineer (LLM Agents & RAG)
Job details
About this role
Role overview
This is a production engineering role focused on building LLM-powered agents that handle real customer work, from answering support questions grounded in client documents to qualifying leads and screening candidates before handing off to a human. Every agent ships to live users, so reliability, evaluation, and ongoing improvement are central rather than optional.
Responsibilities
- Design agents that answer from client data using retrieval, tool use, and clear guardrails. - Connect agents to existing client systems such as help desks, CRMs, and messaging platforms. - Build evaluation sets from real conversations and measure quality before and after every change. - Monitor live agents, review transcripts, and iteratively improve answers. - Explain technical trade-offs to non-technical clients in plain language.
Requirements
- Prior experience shipping LLM-based features to production, with awareness of why some failed. - Proficiency in TypeScript or Python and comfort working with APIs, queues, and databases. - Strong orientation toward evaluation and quality measurement rather than demos alone. - Ability to deliver in weekly increments and present work regularly.
Nice to have
- Hands-on experience with retrieval, embeddings, vector or hybrid search. - Familiarity with Zendesk, Intercom, or the WhatsApp Business Platform.
Benefits and work setup
Fully remote team collaborating with clients across the US, UK, and India, with weekly working software shipped and demonstrated to clients. Source code and infrastructure are handed over to clients at the end of each engagement.