← Back to jobs

Remote job

DevOps / SRE Lead

DevOps Remote-eligible

Job details

Not specified Salary
Remote-eligible Eligibility
Lead Experience
Not specified Employment

About this role

Role overview Lead the reliability discipline for a multi-region AI infrastructure platform, including a GPU cloud, an internal data platform, and customer-facing APIs. The role blends hands-on technical work with cultural leadership, owning service level objectives, incident response, and observability standards across the engineering organisation. It is a remote, full-time lead position focused on turning availability from an aspiration into a measurable, contractual reality.

Responsibilities

- Define and maintain service level objectives and error budgets across the GPU cloud, internal platform, and public APIs, and ensure those budgets actually shape what ships. - Own the observability stack end to end, including metrics, logs, traces, dashboards, and the signal-to-noise quality of alerting. - Run incident management end to end: on-call rotation design, incident command, escalation paths, and blameless post-mortems that produce durable fixes. - Drive reliability improvements to completion, tracking post-mortem actions rather than filing and forgetting them. - Set infrastructure-as-code standards and change-management discipline so risky changes are staged and reversible. - Maintain runbooks and operational documentation, keeping them current as systems evolve, and coach engineers across teams on designing operability in from the start.

Requirements

- Substantial experience running multi-region production systems with genuine availability commitments behind them. - Depth in modern observability tooling such as Prometheus, Grafana, Loki, OpenTelemetry, or comparable equivalents. - Practical experience defining and operating SLOs and error budgets, including the organisational conversations they force. - Strong infrastructure-as-code background, such as Terraform or Pulumi, paired with disciplined code review. - Proven incident-command capability under real pressure, with the temperament that goes with it. - Automation-first instincts backed by working coding ability in Python or Go. - Comfort leading through influence across teams that do not report to you.

Nice to have

- Experience operating GPU or other specialised hardware fleets. - Background in capacity planning and cost optimisation for expensive compute. - Exposure to chaos engineering or systematic resilience testing. - Security or compliance experience relevant to enterprise customers.

Benefits and work setup

- Remote-eligible, full-time position with competitive compensation.

Skills detected in the listing

PythonGoAWSTerraform
Detected Sep 28, 2026
Last verified Not yet verified

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight