Remote job
Staff Software Engineer, Agentic Tools
Job details
About this role
Role overview This role focuses on building a shared AI platform and productivity layer for an engineering organization, designing production agentic workflows, orchestration, and platform harnesses that other teams can extend. The work centers on intent engineering, runtime infrastructure, and observability so that AI-augmented software delivery becomes trusted and broadly adopted. The position is remote-first and may include stand-by, on-call, or off-hours duties for systems owned.
Responsibilities - Design, build, and operate production agentic workflows and the platform harnesses they run on, including orchestration, tool integrations, shared context, and extension points for partner teams. - Practice intent engineering by turning goals into clear specifications, rules, constraints, and acceptance criteria that AI systems can execute reliably. - Build observability and guardrails: tracing, regression detection, human-in-the-loop controls, safe rollout, and operability for agentic systems others depend on. - Own critical pieces of agent runtime, including job isolation, scheduling, execution environments, secrets and access, and cloud production operations. - Design event-driven architectures where appropriate, using queues, streams, webhooks, and async job fan-out to keep agentic systems scalable and loosely coupled. - Drive adoption through playbooks, examples, demos, and enablement, while staying current on models, agent frameworks, and MCP-style tool protocols.
Requirements - Staff-level professional software engineering experience building and operating production distributed systems. - Strong TypeScript and modern JavaScript plus Node, with experience building backend services, integrations, automation, and internal tools. - Professional experience with agent orchestration, tool-calling systems, evaluation or guardrail techniques, and connecting agents to reliable backend systems. - Familiarity with agent frameworks, orchestration layers, and integrating external tools and data sources into LLM-based systems (MCP or equivalent a plus). - Experience with event-driven systems such as queues, streams, pub/sub, or webhooks in production. - Observability fluency across metrics, logs, and traces, with clear communication, documentation, and teaching ability to drive adoption.
Nice to have - Hands-on Kubernetes in production - Terraform or similar infrastructure-as-code for cloud provisioning - Temporal or other workflow-platform experience - Deeper AWS or cloud-native operations fluency (IAM, networking, running production workloads end to end)
Benefits and work setup - Base salary range of $210,000–$245,000 USD, depending on geographic location, experience, and skills - Eligibility to participate in the corporate bonus program - Health, dental, vision, short-term disability, and life insurance - Paid holidays and paid time off - Fertility treatment benefit, 401(k), and equity - Remote-first work environment