Remote job
Senior Site Reliability Engineer (Remote)
Job details
About this role
Role overview
A senior platform reliability engineer is needed to run and evolve the production infrastructure behind a cloud computing portfolio. The role blends deep Kubernetes and Linux engineering with network architecture, observability, and SRE practice across bare-metal, on-prem, and virtualized environments. It is fully remote within EU timezones and suits someone who values autonomy, broad technical scope, and high-impact infrastructure work.
Responsibilities
- Operate and harden Linux infrastructure (Debian/Ubuntu) across bare-metal, virtualized, and on-prem footprints. - Deploy, scale, and maintain Kubernetes clusters end-to-end, including upgrades, node pools, networking, storage, and security hardening. - Design network architecture spanning VLANs, L2/L3 routing, VPNs, and multi-site connectivity. - Build automation for provisioning and operations using Ansible, Bash/Python, and GitOps workflows, including PXE boot, Preseed, and cloud-init. - Stand up and tune observability stacks (Prometheus/Grafana, Loki, ELK, Graylog) and define SLOs/SLIs across physical, virtualization, and application layers. - Lead incident response, coordinate on-call coverage across timezones, author SOPs, and partner with engineering and customer-facing teams on reliability, capacity planning, and architecture decisions.
Requirements
- Expert-level, hands-on experience running Kubernetes in production. - Strong network engineering skills across VLANs, L2/L3 routing, VPNs, and multi-site designs. - Solid Linux administration background on Debian/Ubuntu. - Track record building automation with Ansible, Bash/Python, and Git-based workflows. - Practical experience with observability tooling such as Prometheus, Grafana, ELK, Loki, or Graylog. - Familiarity with virtualization platforms (OpenStack, Proxmox, VMware), bare-metal provisioning with MAAS, plus distributed systems fundamentals and incident or on-call practices.
Nice to have
- Service mesh experience (Istio, Linkerd) or advanced CNI implementations. - Cloudflare APIs, DNS automation, or tunnel configuration knowledge. - GPU infrastructure, node preparation, or resource scheduling exposure. - Security best practices including RBAC, firewalls, and network policies. - IT asset management or license tracking workflows. - Coordinating reliability practices across distributed, multi-timezone teams.
Benefits and work setup
- 100% remote role with flexible hours within an EU timezone window. - High-impact position with autonomy and ownership over platform direction. - International, collaborative engineering team. - Modern stack with a strong emphasis on reliability engineering and automation.