← Back to jobs

Remote job

Site Reliability Tech Lead, AI-DNA

DevOps Remote, async-first, global; US-morning overlap required

Job details

Not specified Salary
Remote, async-first, global; US-morning overlap required Eligibility
Lead Experience
Not specified Employment

About this role

Role overview

Lead reliability engineering for a product family supporting high-stakes enterprise workloads. The role combines hands-on site reliability engineering, incident command, and the design of AI agents that can investigate incidents, validate changes, draft root-cause analyses, and apply approved remediation. You will set the technical quality bar for a small, senior team without moving into people management.

Responsibilities

- Own availability, service-level performance, time to acknowledge, and mean time to recovery for the assigned product area. - Serve as technical owner and incident commander during major outages, while partnering with executive stakeholders on customer-impacting events. - Design and decompose an AI-agent operating system for pre-triage, change validation, permanent-fix tracking, root-cause analysis, and policy-controlled auto-healing. - Establish safe production pathways for deployments, configuration changes, and cost-optimization runbooks, including tested rollback procedures. - Review complex changes and quality-check incident analyses before they reach production or are shared beyond the team. - Turn recurring incidents into durable automation, better safeguards, and measurable week-over-week reliability improvements.

Requirements

- Deep, battle-tested experience in site reliability engineering, production operations, and complex incident response. - Strong ability to design maintainable automation and agent-based operational workflows, with careful attention to failure modes and guardrails. - Experience owning reliability for distributed, customer-facing systems at meaningful scale. - Ability to work directly in agent or automation code while making sound operational and architectural decisions. - Comfort operating in a remote, async-first environment with rapid decisions and high individual ownership. - A demonstrated record of depth in a difficult technical area, such as original technical writing, talks, or open-source work.

Nice to have

- Experience with AWS plus exposure to Azure. - Familiarity with observability and incident-response tools such as Grafana, Prometheus, PagerDuty, OpsGenie, or Datadog. - Public or open-source contributions focused on agentic SRE, AIOps, or related operational automation.

Benefits and work setup

Remote and globally distributed, with expected overlap during US mornings. The team works asynchronously, operates with a small senior footprint, and treats AI agents as part of the operating model.

Skills detected in the listing

AWSAzure
Detected Sep 19, 2026
Last verified Sep 19, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight