Remote job
Site Reliability Engineer, Infrastructure Platforms — UK (Intermediate to Senior Staff)
Job details
About this role
Role overview This intermediate-to-senior staff site reliability engineering role supports the infrastructure platforms that a large software organization depends on for building, testing, and shipping code. The position is remote within the United Kingdom and covers both hands-on operational work and platform-level engineering for the systems underneath the product. It spans reliability, automation, and capacity concerns for shared infrastructure.
Responsibilities - Operate and improve the reliability of core infrastructure platforms, including compute, storage, and supporting services. - Build and maintain automation, monitoring, and incident-response tooling. - Participate in on-call rotations, lead post-incident reviews, and drive follow-up remediation work. - Partner with platform engineering teams to evolve infrastructure-as-code, deployment, and capacity-planning practices. - Contribute to design reviews for new platform features with a focus on operability and scale.
Requirements - Substantial experience as an SRE, platform engineer, or infrastructure software engineer, with progressive responsibility. - Strong knowledge of Linux, Kubernetes, infrastructure-as-code tooling, and at least one major cloud provider. - Proficiency with monitoring, logging, and tracing systems, and a data-driven approach to incident analysis. - Comfort writing production code, reviewing infrastructure changes, and mentoring peers.
Nice to have - Experience operating multi-region or high-traffic SaaS platforms. - Familiarity with CI/CD systems, networking, and security baselines for cloud environments. - Exposure to capacity modeling, cost optimization, or internal developer platform programs.