Remote job
Site Reliability Tech Lead, AI-DNA
Job details
About this role
Role overview
Lead reliability engineering for a product family supporting high-stakes enterprise workloads. The role combines hands-on site reliability engineering, incident command, and the design of AI agents that can investigate incidents, validate changes, draft root-cause analyses, and apply approved remediation. You will set the technical quality bar for a small, senior team without moving into people management.
Responsibilities
- Own availability, service-level performance, time to acknowledge, and mean time to recovery for the assigned product area. - Serve as technical owner and incident commander during major outages, while partnering with executive stakeholders on customer-impacting events. - Design and decompose an AI-agent operating system for pre-triage, change validation, permanent-fix tracking, root-cause analysis, and policy-controlled auto-healing. - Establish safe production pathways for deployments, configuration changes, and cost-optimization runbooks, including tested rollback procedures. - Review complex changes and quality-check incident analyses before they reach production or are shared beyond the team. - Turn recurring incidents into durable automation, better safeguards, and measurable week-over-week reliability improvements.
Requirements
- Deep, battle-tested experience in site reliability engineering, production operations, and complex incident response. - Strong ability to design maintainable automation and agent-based operational workflows, with careful attention to failure modes and guardrails. - Experience owning reliability for distributed, customer-facing systems at meaningful scale. - Ability to work directly in agent or automation code while making sound operational and architectural decisions. - Comfort operating in a remote, async-first environment with rapid decisions and high individual ownership. - A demonstrated record of depth in a difficult technical area, such as original technical writing, talks, or open-source work.
Nice to have
- Experience with AWS plus exposure to Azure. - Familiarity with observability and incident-response tools such as Grafana, Prometheus, PagerDuty, OpsGenie, or Datadog. - Public or open-source contributions focused on agentic SRE, AIOps, or related operational automation.
Benefits and work setup
Remote and globally distributed, with expected overlap during US mornings. The team works asynchronously, operates with a small senior footprint, and treats AI agents as part of the operating model.