Remote job
Site Reliability Engineer, Provider Operations
Job details
About this role
Role overview
Own the operational health and reliability of external AI inference providers and their endpoints. The role focuses on detecting performance and quality problems early, shifting traffic away from unhealthy services, and building the monitoring and automation needed to keep customer-facing systems dependable at high request volumes.
Responsibilities - Build monitoring for provider and endpoint latency, throughput, errors, availability, and output correctness; define service objectives and actionable alerts. - Improve detection of degraded endpoints and coordinate automatic traffic failover with routing teams. - Lead incident response, including triage, mitigation, communication with providers, postmortems, and follow-up actions. - Turn operational telemetry into provider scorecards and service-level reporting, and act as a technical escalation contact. - Create canaries and evaluations that identify quality regressions such as broken tool calls, truncated streams, or usage-reporting mismatches. - Replace manual provider operations with safe, auditable tools, and build load-testing tools to prepare endpoints for launches.
Requirements - At least 4 years in site reliability, production engineering, or infrastructure roles operating high-traffic customer-facing systems. - Strong observability experience across metrics, tracing, logs, service objectives, error budgets, and alerting. - Software engineering ability in TypeScript and/or Python, with a preference for building tools to reduce manual work. - Understanding of distributed-system failure modes such as timeouts, retries, backpressure, partial outages, and noisy neighbors. - Calm incident leadership and clear communication with external partners under pressure. - Understanding of AI inference serving, or the ability and interest to learn its streaming, tool-calling, caching, and performance tradeoffs.
Nice to have
Experience with an inference provider, model lab, GPU cloud, API gateway, or CDN; routing or load balancing; or evaluation and synthetic monitoring for machine learning systems. Familiarity with TypeScript, Cloudflare Workers, Postgres, ClickHouse, GCP, or Vercel is also useful.