← Back to jobs

Remote job

Site Reliability Engineer, Provider Operations

DevOps Full-time Permanent United States

Job details

Not specified Salary
United States Eligibility
Not specified Experience
Full-time Employment

About this role

Role overview

Own the operational health and reliability of external AI inference providers and their endpoints. The role focuses on detecting performance and quality problems early, shifting traffic away from unhealthy services, and building the monitoring and automation needed to keep customer-facing systems dependable at high request volumes.

Responsibilities - Build monitoring for provider and endpoint latency, throughput, errors, availability, and output correctness; define service objectives and actionable alerts. - Improve detection of degraded endpoints and coordinate automatic traffic failover with routing teams. - Lead incident response, including triage, mitigation, communication with providers, postmortems, and follow-up actions. - Turn operational telemetry into provider scorecards and service-level reporting, and act as a technical escalation contact. - Create canaries and evaluations that identify quality regressions such as broken tool calls, truncated streams, or usage-reporting mismatches. - Replace manual provider operations with safe, auditable tools, and build load-testing tools to prepare endpoints for launches.

Requirements - At least 4 years in site reliability, production engineering, or infrastructure roles operating high-traffic customer-facing systems. - Strong observability experience across metrics, tracing, logs, service objectives, error budgets, and alerting. - Software engineering ability in TypeScript and/or Python, with a preference for building tools to reduce manual work. - Understanding of distributed-system failure modes such as timeouts, retries, backpressure, partial outages, and noisy neighbors. - Calm incident leadership and clear communication with external partners under pressure. - Understanding of AI inference serving, or the ability and interest to learn its streaming, tool-calling, caching, and performance tradeoffs.

Nice to have

Experience with an inference provider, model lab, GPU cloud, API gateway, or CDN; routing or load balancing; or evaluation and synthetic monitoring for machine learning systems. Familiarity with TypeScript, Cloudflare Workers, Postgres, ClickHouse, GCP, or Vercel is also useful.

Skills detected in the listing

TypeScriptPythonGoPostgreSQLGCPLLM
Detected Oct 9, 2026
Last verified Oct 9, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight