Remote job
Staff Engineer, Reliability & Routing
Job details
About this role
Role overview
This senior individual-contributor position sets the technical direction for an API gateway that routes customer traffic across multiple upstream providers. The role blends hands-on engineering with architectural ownership, focused on latency-sensitive routing, fast failover, and deep observability at the edge of the network.
Responsibilities
- Own route resolution logic, including weighted and priority strategies, per-target retries, circuit breaking, and connection health checks - Define and document what the gateway guarantees and explicitly what it declines to do (for example, never re-routing after the first streamed byte) - Build failure drills and load tests that prove failover behavior before customers encounter it - Design telemetry that lets an operator quickly identify which provider failed, for which tenant, and since when - Review designs across the gateway and raise the standard for correctness and operability
Requirements
- 8+ years of backend or infrastructure engineering, including responsibility for a latency-sensitive production service - Deep practical experience with retry and backoff policy, timeouts, load shedding, and failure isolation - Hands-on incident response experience and a record of writing blameless postmortems - Strong proficiency in TypeScript or another typed language, with willingness to work in TypeScript day to day - A track record of technical leadership through written design documents
Nice to have
- Experience building an API gateway, proxy, or service mesh - Familiarity with OpenTelemetry tracing and semantic conventions - Operating on an edge or serverless runtime where CPU time, not wall time, is the budget