Remote job
Operational Engineer
Job details
About this role
Role overview This role focuses on making engineering services reliable, secure, measurable, and easy to operate within a GPU cloud environment built for AI workloads. The position sits at the intersection of operations engineering and platform engineering, building shared practices, tooling, reporting, and automations that other service teams depend on. The work is hands-on, fast-paced, and centred on turning recurring operational problems into durable improvements.
Responsibilities - Onboard services to operational tooling covering alerting, on-call rotations, dashboards, runbooks, ticketing, and service-catalog records. - Build and maintain integrations, dashboards, reports, and automations across monitoring, alerting, ticketing, internal developer portals, and cloud billing systems. - Improve operational readiness by identifying gaps in ownership, alerting, documentation, recovery procedures, and service dependencies. - Act as incident commander during major events, coordinating response, communications, timelines, post-incident reviews, and follow-up actions. - Make routine changes safer and more repeatable through clear processes and automation, and reduce manual toil across teams. - Contribute to service-health, SLA, cost, patching, and operational-risk reporting, keeping ownership and dependency data accurate.
Requirements - 2-5 years of experience in operations engineering, SRE, cloud infrastructure, platform engineering, or a closely related role. - Hands-on experience operating or supporting production services in a cloud or infrastructure context. - Working knowledge of monitoring, alerting, on-call practices, runbooks, and engineering work management tools. - Practical experience with at least some of Grafana, PagerDuty, Jira, Backstage, public cloud platforms, dashboards, integrations, or workflow automation. - Ability to analyse operational data, spot gaps, and drive actions through to closure.
Nice to have - Familiarity with GPU clusters, AI workloads, or high-performance compute environments. - Exposure to FinOps or cloud cost reporting and tooling.