Remote job
Agentic Systems Engineer
Job details
About this role
Role overview This is the core build seat for agentic systems, focused on designing how agents are decomposed, how they call tools, how they hand off work, how they remember, and how to detect when they fail. The role spans the Claude and OpenAI APIs, open weight models, and the Model Context Protocol, with substantial time spent on the parts that are not the model: the orchestration runtime, the evaluation harness, and the cost and latency controls. The goal is turning a working demo into a product that holds up under real traffic.
Responsibilities - Decompose problems into single agent, multi agent, or plain workflow designs, and defend the choice with measurements - Build tool layers with typed schemas, bounded side effects, and code-enforced permission models, exposing them over MCP where multiple agents consume them - Design memory and context management, including retrieval, summarization, prompt caching, and context budgets that hold up at real conversation lengths - Implement orchestration in TypeScript or Python, covering planning loops, handoffs, parallel tool execution, cancellation, and durable state across long-running tasks - Engineer cost and latency with model routing between frontier and small models, caching, streaming, and batching, with per-task budgets that alarm when exceeded - Build evaluation suites from real traces, run them on every prompt or model change in CI, and instrument traces so non-engineers can read what an agent did and why
Requirements - Five or more years of production software engineering in TypeScript or Python - At least one agentic or LLM tool calling system built and operated for real users - Deep familiarity with a frontier model API such as Claude, OpenAI, or Gemini, including tool use, structured output, streaming, and caching behavior - A working evaluation harness used to drive decisions, not just produce charts - Strong systems instincts around idempotency, timeouts, retries, and state - Ability to explain what an agent is and is not good at without overselling
Nice to have - Experience with MCP servers or clients, the Claude Agent SDK, or comparable orchestration frameworks - Running open weight models locally or self-hosted with vLLM or MLX, and understanding how tool calling degrades on smaller models - Background in distributed systems, workflow engines, or durable execution