Remote job
Senior Operational Engineer
Job details
About this role
Role overview A GPU cloud provider built for AI workloads is hiring a Senior Operational Engineer to lead cross-service improvements that make its engineering services more reliable, secure, controlled, and cost-aware. The role designs and delivers shared operational capabilities such as readiness standards, canary checks, runbooks, reporting, alerting workflows, and automation, while partnering with service teams to address systemic operational risk. It is a hands-on technical position with significant influence across the wider engineering organization.
Responsibilities - Take ownership of designing and delivering shared operational capabilities, including readiness checks, canaries, runbooks, service-health reporting, and operational automation. - Establish practical standards for service ownership, on-call readiness, alerts, dashboards, recovery procedures, and operational evidence. - Partner with service teams to identify and resolve recurring operational issues and cross-team blockers. - Design and improve integrations and workflows across Grafana, PagerDuty, Jira, Backstage, public-cloud platforms, and reporting systems. - Turn incident findings, change failures, and near misses into lasting engineering improvements, and act as incident commander leading technical analysis and corrective-action planning. - Build safe, observable, auditable automations that reduce manual work, and automate operational reporting on SLA performance, health indicators, cost trends, patching status, and action closure. - Mentor engineers, review operational designs, contribute to continuity testing, and run focused failure-and-recovery experiments.
Requirements - Six to ten years of experience in operations engineering, site reliability engineering, cloud infrastructure, platform engineering, or a comparable production-focused engineering role. - Experience designing and operating shared capabilities for monitoring, alerting, on-call, runbooks, service readiness, and operational reporting. - Hands-on experience building integrations and automations across Grafana, PagerDuty, Jira, Backstage, and at least one public cloud platform. - Track record of using incident, change, service-health, continuity, patching, or cost data to drive operational improvement. - Ability to lead cross-service work, set practical standards, and support teams through incidents as a senior technical leader.