Remote job
Senior/Staff Software Engineer (Platform and Execution Model)
Job details
About this role
Role overview
This Senior/Staff Software Engineer will own critical pieces of an enterprise AI platform's shared execution substrate that powers deployments in regulated environments. The work centers on distributed systems primitives that connect long-lived workflows, agents, tools, and product surfaces, with a heavy emphasis on reliability, scalability, and correctness under failure. Staff-level scope expands the role to leading architecture across teams, establishing engineering standards, and mentoring other engineers.
Responsibilities
- Design and own core execution model components, including the state machine, lifecycle, resource model, and failure semantics. - Build distributed systems that behave predictably across retries, restarts, partial failures, and concurrent execution. - Develop and evolve platform APIs and SDKs that connect workflows, agents, tools, and product surfaces, including versioning and compatibility strategy. - Engineer reliability primitives such as concurrency controls, rate limits, backpressure, sharding, partitioning, and workload isolation. - Embed security and governance into the core through RBAC/ABAC, policy enforcement, and fine-grained audit and lineage. - Deliver observability via distributed tracing, structured logs, metrics, and evaluation hooks, building an explainable trail of agent actions. - Drive quality through design reviews, test strategy, performance baselines, SLOs, incident response, and postmortems; mentor engineers and contribute to engineering patterns across the platform.
Requirements
- 8+ years of experience building distributed or platform systems, including significant time defining architecture across teams or domains. - 4+ years owning mission-critical runtimes or workflow/orchestration systems. - Deep expertise in durable execution patterns such as state machines, event sourcing, saga and compensation logic, idempotency, and exactly- or at-least-once delivery semantics. - Proven track record with production security and governance, including authentication, RBAC, audit, and policy enforcement. - Hands-on experience with observability tooling (Grafana or equivalent), including trace correlation across asynchronous boundaries. - Strong systems design across storage, queues, schedulers, and evented architectures, with comfort tuning performance under load. - Proficiency in a modern language such as Go, Rust, Java, or TypeScript, plus cloud-native stacks including containers, CI/CD, and infrastructure as code. - Comfort operating in regulated or high-assurance environments, with a bias toward correctness, clarity, and thorough documentation.
Nice to have
- Prior work on workflow engines such as Temporal, Cadence, AWS Step Functions, Argo, or Airflow, or on serverless runtimes. - Experience with policy engines (for example OPA), secrets and KMS, or data-handling controls for PII or PHI. - Familiarity with ML/LLM evaluation frameworks, tool and plugin architectures, or embedding model governance into execution flows. - Government or healthcare background, including HIPAA exposure, audit readiness, and multi-tenant isolation.
Benefits and work setup
- Comprehensive medical, dental, and vision insurance fully paid by the employer for employees and their families. - 14 weeks of paid maternity and paternity leave at normal pay. - Unlimited PTO subject to management approval. - Optional 401(k), FSA, and equity incentives. - Mental health benefits and access to cost-effective GLP-1 care programs. - Professional development budget and a career-track opportunity with potential for rapid advancement. - Some travel to customer sites may be required.