Remote job
Senior/Lead Fullstack Developer
Job details
About this role
Role overview A remote Senior or Lead Fullstack Developer is needed to build and optimize an advanced retrieval-augmented generation (RAG) platform. The work centers on improving how context is preserved during information retrieval, with the goal of helping customers extract more value from their knowledge bases across diverse domains. The team works with modern embedding models, vector databases, and large language models to push retrieval accuracy forward.
Responsibilities - Implement and optimize contextual embedding and document preprocessing pipelines. - Design efficient document chunking and context generation workflows. - Build reranking systems that further improve retrieval accuracy. - Develop evaluation frameworks that measure retrieval performance across different domains. - Create cost-effective solutions using prompt caching and related optimization techniques. - Experiment with different embedding models to identify optimal configurations.
Requirements - Experience working with embedding models such as Gemini, Voyage, or similar systems. - Strong understanding of vector databases and similarity search. - Hands-on prompt engineering experience for context generation. - Familiarity with document chunking strategies and preprocessing techniques. - Understanding of reranking systems and how to implement them. - Background in measuring retrieval accuracy, including recall metrics. - Programming experience in Python or a comparable language. - Experience with large-scale data processing.
Nice to have - Background in NLP, information retrieval, or computational linguistics. - Experience architecting and shipping RAG systems in production. - Knowledge of LLM prompt caching and cost optimization techniques. - Familiarity with domain-specific retrieval challenges such as code or scientific papers. - Experience balancing performance improvements against latency and cost. - Background working with major enterprise LLMs. - Understanding of AWS or comparable cloud infrastructure for LLM applications.