Remote job
AI Platform Engineer
Job details
About this role
Role overview A distributed, open-source-first software company is standing up a small central team to build the internal "execution layer" for AI agents that carry real operational load across the business. As the sole platform engineer on that team, you will design and ship the runtime, queue, evaluation harness, and governance controls that make agentic work trustworthy, observable, and structurally bounded.
Responsibilities
- Build and ship the production agent platform: an event-triggered queue, a model-agnostic headless runtime, durable state that survives restarts, a human review gate, atomic rollback, and complete run logging in the warehouse. - Build the evaluation layer: golden suites with behavioral assertions, a judge rubric, safety cases that must pass every run, and a CI gate that blocks regressions from merging. - Register and maintain the agent portfolio across executive, team-lead, and individual-contributor workflows, plus a meta-layer that observes the platform and improves it. - Enforce governance in code through risk tiers, least-privilege credentials per agent, tool-permission gates, an audit log, and autonomy classes where dangerous actions have no code path. - Define how the system contacts people: hard per-person interruption budgets, message bundling, and quiet hours, so notifications spend trust rather than burn it. - Instrument the system so it measures and publishes the operational work it has absorbed, using that data to argue for or against expansion.
Requirements
- Track record of shipping production LLM agent systems that other people depended on, with operational history, real users, and at least one incident you can discuss. - Experience designing evaluations rather than spot checks: golden sets, behavioral assertions, judge rubrics, pass thresholds, and CI gating. - Deep API work against the systems where work actually lives, including authoring MCP servers and knowing their specific failure modes. - End-to-end infrastructure ownership in Python on a major cloud, with a cloud warehouse and infrastructure-as-code, including provisioning, deployment, monitoring, and rollback. - Strong taste for how software contacts humans, treating every notification as a limited trust expenditure. - Proficiency with agentic tools on real work in files and repositories, and the ability to describe your own setup mechanically, including triggers, permissions, approvals, and logging.
Nice to have
- Public work in the space: an open-source agent framework, an MCP server, an evaluation harness, or writing on agent reliability cited by other practitioners. - LLM observability and cost instrumentation, including tracing agent runs, attributing spend per run, and building the queries that prove the platform's return.
Benefits and work setup
- Fully remote, async-first work with global hiring and a co-working allowance usable anywhere. - Equity ownership for every team member, plus a tech allowance for personal equipment. - Health coverage for employees and partial coverage for dependents, regardless of location. - Annual company off-sites, flexible scheduling, and an annual professional development allowance.