Remote job
Senior/Staff Software Engineer (Platform and Execution Model)
Job details
About this role
Role overview
This Senior/Staff Software Engineer will own critical parts of an enterprise AI platform's shared execution substrate that powers deployments in regulated industries. The position focuses on distributed systems and platform primitives that connect long-lived workflows, agents, tools, and product surfaces, prioritizing reliability, scalability, and correctness under failure. Staff-level scope broadens into leading architecture across teams, setting engineering standards, and mentoring engineers.
Responsibilities
- Design and own critical components of the core execution model, spanning the state machine, lifecycle, resource model, and failure semantics. - Build distributed systems that behave predictably across retries, restarts, partial failures, and concurrent execution. - Develop platform APIs and SDKs connecting workflows, agents, tools, and product surfaces, including the versioning and compatibility strategy. - Engineer reliability at scale through concurrency controls, rate limits, backpressure, sharding and partitioning, and workload isolation. - Build security and governance into the core with RBAC/ABAC, policy enforcement, and fine-grained audit and lineage. - Deliver observability via distributed tracing, structured logs, metrics, and evaluation hooks to maintain an explainable trail of agent actions. - Own quality through design reviews, test strategy (unit, property, chaos), performance baselines, SLOs, incident response, and postmortems; mentor engineers and shape platform-wide engineering patterns.
Requirements
- 8+ years building distributed or platform systems, including meaningful experience defining architecture across teams or domains. - 4+ years owning mission-critical runtimes or workflow/orchestration systems. - Deep expertise in durable execution patterns such as state machines, event sourcing, saga or compensation logic, idempotency, and exactly- or at-least-once semantics. - Track record with production security and governance, including authentication, RBAC, audit, and policy enforcement. - Hands-on experience with observability stacks (Grafana or equivalent), including tracing across asynchronous boundaries. - Strong systems design across storage, queues, schedulers, and evented architectures; comfortable tuning performance under load. - Excellence in a modern language such as Go, Rust, Java, or TypeScript, plus cloud-native stacks (containers, CI/CD, infrastructure as code). - Comfortable in regulated or high-assurance environments, with a bias toward correctness, clarity, and documentation.
Nice to have
- Prior work on workflow engines like Temporal, Cadence, AWS Step Functions, Argo, or Airflow, or on serverless runtimes. - Experience with policy engines such as OPA, secrets and KMS handling, or PII/PHI data controls. - Familiarity with ML/LLM evaluation frameworks, tool and plugin architectures, or embedding model governance into execution flows. - Government or healthcare exposure (HIPAA, audit readiness) and multi-tenant isolation work.
Benefits and work setup
- Salary range of $180,000–$240,000 depending on experience and skills. - Comprehensive medical, dental, and vision coverage fully paid by the employer for employees and their families. - 14 weeks of paid maternity and paternity leave at normal pay. - Unlimited PTO with management approval. - Optional 401(k), FSA, and equity incentives. - Mental health benefits and access to cost-effective GLP-1 care programs. - Professional development support and a career-track opportunity with potential for rapid advancement as the firm grows. - Some travel to customer sites may be required.