Remote job
Senior Site Reliability Engineer / Kubernetes (Remote)
Job details
About this role
Role overview
A senior platform role focused on operating and scaling production Kubernetes infrastructure for a cloud computing organization. The position combines hands-on SRE work with architectural ownership across bare-metal, virtualized, and on-prem environments, with a strong emphasis on reliability, automation, and networking. It is a fully remote position aligned to European time zones, offering significant autonomy and regular collaboration with cross-functional engineering teams.
Responsibilities
- Operate and maintain Debian/Ubuntu-based Linux infrastructure, including the full Kubernetes cluster lifecycle: upgrades, node pools, networking, storage, and security hardening. - Design and maintain network architecture spanning VLANs, L2/L3 routing, VPNs, and multi-site connectivity. - Build automation for provisioning and operations using Ansible, Bash/Python, and GitOps workflows, including bare-metal deployment pipelines (PXE boot, Preseed, cloud-init). - Deploy and operate observability stacks such as Prometheus, Grafana, Loki, ELK, or Graylog, and define SLOs/SLIs across physical, virtualization, and service layers. - Lead incident response, manage on-call rotations, and author Standard Operating Procedures for repeatable operations and maintenance. - Manage virtualization and orchestration layers (OpenStack, Proxmox, VMware), coordinate physical hardware maintenance, and collaborate with development and customer-facing teams on capacity planning and architecture.
Requirements
- Expert-level, hands-on experience operating Kubernetes in production environments. - Strong network engineering skills across VLANs, L2/L3 routing, VPNs, and multi-site connectivity. - Solid Linux systems administration background on Debian/Ubuntu, with a deep understanding of distributed systems and container orchestration. - Experience building automation with Ansible, Bash/Python, and Git-based workflows, plus familiarity with observability tooling (Prometheus, Grafana, ELK, Loki, or Graylog). - Background with virtualization platforms (OpenStack, Proxmox, VMware) and bare-metal provisioning using MAAS. - Demonstrated incident response experience, ability to author SOPs and on-call processes, and comfort working autonomously in a fast-paced, engineering-driven environment.
Nice to have
- Service mesh experience (Istio, Linkerd) or advanced CNI implementations. - Knowledge of Cloudflare APIs, DNS automation, or tunnel configurations. - Experience with GPU infrastructure, node preparation, or resource scheduling. - Familiarity with security best practices such as RBAC, firewalls, and network policies. - Exposure to IT asset management or license tracking workflows. - Background establishing reliability practices and SRE frameworks in growing organizations, including coordination across distributed, multi-timezone teams.
Benefits and work setup
- Fully remote with flexible hours, aligned to CET ±2h. - High-impact role with autonomy and ownership of platform decisions. - Collaborative, international engineering team. - Modern technology stack with a strong focus on reliability and automation.