Remote job
Senior Site Reliability Engineer, AI-DNA
Job details
About this role
Role overview
Senior site reliability role centered on operating a high-stakes, multi-tenant SaaS platform and turning operational expertise into AI-powered automation. You will respond to incidents, execute safe production changes, and build agents and runbooks that progressively reduce repetitive manual intervention while preserving enterprise reliability standards.
Responsibilities
- Serve as a first responder for production incidents, triaging customer impact and escalating when the blast radius requires it. - Execute deployments, configuration updates, and cost-optimization changes with validation, quality gates, and rollback plans. - Build, operate, and maintain agents for pre-triage, change validation, root-cause analysis, permanent-fix follow-up, and auto-healing. - Create human and AI runbooks that turn recurring operational problems into repeatable, increasingly automated solutions. - Generalize one-off incident fixes so similar failures can be handled by agents in the future. - Identify operational gaps independently and improve the platform, tooling, and response process without waiting for task assignment.
Requirements
- Senior-level SRE or production operations experience with the ability to run incident response independently. - Strong AWS experience and practical knowledge of reliability, observability, deployment, and incident-management practices. - Ability to design safe operational automation with clear policy boundaries and rollback behavior. - Comfortable working with AI agents, runbooks, logs, code paths, and historical incident data. - Ownership mindset, sound judgment under pressure, and willingness to work in an asynchronous environment. - Clear country of residence, as requested in the source requirements.
Nice to have
- Experience with Grafana, Prometheus, OpsGenie, PagerDuty, Datadog, Azure, or multi-tenant B2B SaaS. - Evidence of building or sharing agentic SRE or AIOps tools through open source, talks, writing, or shipped systems. - Deep expertise in a difficult technical area beyond mainstream role responsibilities.
Benefits and work setup
- Remote role with enterprise-scale customers and a fast, startup-style operating cadence. - Resources are available for additional compute, models, or tools when they strengthen the operational harness. - Opportunity to gain hands-on experience building agent-driven SRE systems and self-running operational runbooks.