Remote job
Site Reliability Engineer
Job details
About this role
Role overview This position focuses on advancing next-generation AI agent and retrieval systems in a production setting. The work blends research and engineering, building adaptive Agentic RAG pipelines, agent orchestration runtimes, and evaluation frameworks that turn real-world signals into measurable improvements. It suits someone who moves fluidly between research questions and shipped systems at the intersection of large language models, retrieval, and autonomous workflows.
Responsibilities - Design and operate retrieval pipelines that go beyond static retrieve-once patterns, including adaptive, self-correcting, and multi-hop workflows - Architect Agentic RAG systems with dynamic retrieval control, query decomposition, and iterative retrieve-reflect-refine loops - Collaborate on model-capability-driven work spanning context management, long-term memory, subagent and multi-agent architectures, and real-world task execution - Propose and build benchmarks and evaluation methodologies for agent and retrieval domains, including datasets, annotations, and quality metrics like groundedness, latency, and task success - Use multi-channel user feedback and task data as primary research signals to design experiments and datasets that continuously improve agent and retrieval performance
Requirements - 2-8+ years of hands-on experience with LLMs, RAG, and AI agent systems in production - End-to-end retrieval pipeline experience, including embedding models, vector stores, hybrid search, reranking, chunking strategy, text cleaning, and multimodal parsing - Practical experience with patterns such as Self-RAG, Corrective RAG, adaptive retrieval, and multi-hop decomposition - Hands-on work with agent orchestration runtimes, including session recovery, sandbox isolation, middleware systems, multi-tenant runtime, plan/execute loops, and retrieval-grounded tool calling - Deep familiarity with LLM and agent fundamentals such as KV cache, agent loops, tool use, planning, MCP, memory, and multi-agent systems, plus prompt and context engineering - Independent research ability: analyze ambiguous problems from first principles, generate original ideas, and iterate quickly from prototype to experiment
Nice to have - A track record as a power user of agent products, with informed judgment about model behavior and failure modes - Strong learning velocity across unfamiliar languages, frameworks, and domains, using AI-assisted workflows to ship quickly
Benefits and work setup - Work-from-home arrangement, with specifics that may vary by team - Competitive salary and company benefits - Opportunities for career growth and continuous learning in a results-driven environment