Remote job
DevOps / SRE Lead
Job details
About this role
Role overview Lead the reliability discipline for a multi-region AI infrastructure platform, including a GPU cloud, an internal data platform, and customer-facing APIs. The role blends hands-on technical work with cultural leadership, owning service level objectives, incident response, and observability standards across the engineering organisation. It is a remote, full-time lead position focused on turning availability from an aspiration into a measurable, contractual reality.
Responsibilities
- Define and maintain service level objectives and error budgets across the GPU cloud, internal platform, and public APIs, and ensure those budgets actually shape what ships. - Own the observability stack end to end, including metrics, logs, traces, dashboards, and the signal-to-noise quality of alerting. - Run incident management end to end: on-call rotation design, incident command, escalation paths, and blameless post-mortems that produce durable fixes. - Drive reliability improvements to completion, tracking post-mortem actions rather than filing and forgetting them. - Set infrastructure-as-code standards and change-management discipline so risky changes are staged and reversible. - Maintain runbooks and operational documentation, keeping them current as systems evolve, and coach engineers across teams on designing operability in from the start.
Requirements
- Substantial experience running multi-region production systems with genuine availability commitments behind them. - Depth in modern observability tooling such as Prometheus, Grafana, Loki, OpenTelemetry, or comparable equivalents. - Practical experience defining and operating SLOs and error budgets, including the organisational conversations they force. - Strong infrastructure-as-code background, such as Terraform or Pulumi, paired with disciplined code review. - Proven incident-command capability under real pressure, with the temperament that goes with it. - Automation-first instincts backed by working coding ability in Python or Go. - Comfort leading through influence across teams that do not report to you.
Nice to have
- Experience operating GPU or other specialised hardware fleets. - Background in capacity planning and cost optimisation for expensive compute. - Exposure to chaos engineering or systematic resilience testing. - Security or compliance experience relevant to enterprise customers.
Benefits and work setup
- Remote-eligible, full-time position with competitive compensation.