Remote job
AI Researcher - Bolter
Job details
About this role
Role overview
An applied AI researcher is needed to make autonomous agents that carry out real work trustworthy, capable, and useful. The work sits squarely between research and product — designing experiments, running evaluations, and turning findings into shipped decisions rather than papers. The role joins a lean founding team working directly with the GM and a product engineer, supported by specialized AI agents, during the platform's validation stage.
Responsibilities
- Design and run experiments that test how the platform's agents behave — reliability, context retention, multi-step task completion, and failure modes. - Build and own the evaluation layer, designing evals that measure whether agents actually complete work, not just whether they sound plausible. - Track the state of the art in LLM agents, tool use, and reliability, and translate those findings into recommendations for what to build next. - Produce clear, actionable recommendations the engineering team can ship directly into the product. - Prototype promising research ideas into working features and hand off what proves useful to engineering. - Shape the research roadmap as the platform grows from the validation stage into broader rollout.
Requirements
- A track record of applied AI or machine learning research, ideally in LLM application design, agentic systems, evaluations, or production AI reliability. - Strong experimental design and analysis skills, with the ability to test a hypothesis rigorously and read results honestly. - Solid engineering fundamentals, with the ability to prototype your own experiments rather than only directing others. - Deep familiarity with large language models, tool use, context management, and the failure modes of agentic systems. - Strong written communication, since research that is not understood and acted on does not help. - Comfort with ambiguity, pragmatism about evidence, and the ability to use AI agents as a force multiplier on a small team.
Nice to have
- Experience with LLM evaluations, hallucination mitigation, or production AI reliability at scale. - Experience building or evaluating agentic systems, tool use, or autonomous workflows. - A public track record such as papers, open source, blog posts, or side projects. - Background in product-led research where the output is a shipped feature rather than only a finding.