Remote job
AI Engineer, Retrieval and Grounding
Job details
About this role
Role overview
Own the retrieval and confidence systems that determine what an AI assistant knows and when it should answer. The work is directly tied to response quality: grounding replies in approved customer content, identifying uncertainty, reducing unsupported claims, and proving improvements with evidence from real conversations.
Responsibilities
- Build and refine retrieval pipelines that ground responses in each customer’s approved content. - Define confidence scoring and the thresholds that should trigger human handoff. - Create evaluation harnesses and datasets based on real customer conversations. - Make hallucinations, unsupported claims, and other failure modes visible and measurable. - Balance response quality with inference latency and operating cost as usage increases. - Assess system changes using data rather than subjective impressions.
Requirements
- Production experience shipping LLM features, beyond prototypes or notebook experiments. - Practical knowledge of embeddings, vector search, chunking, and reranking techniques. - Strong Python or TypeScript skills and the ability to reason about the complete system. - Experience building evaluation pipelines and defending quality conclusions with data. - A skeptical, safety-conscious approach to model output in customer-facing experiences.
Nice to have
- Experience with frontier-model APIs such as the Claude API or comparable services. - A background in search relevance or information retrieval. - Experience with multilingual retrieval, particularly for Arabic, French, Spanish, or Chinese content.