Remote job
Senior Site Reliability Engineer (MAAS)
Job details
About this role
Role overview
A senior infrastructure engineer is needed to run and grow a distributed cloud platform spanning bare-metal GPU nodes, Kubernetes clusters, virtualization layers, and multi-site networking. The role is fully remote and expects end-to-end ownership from the hardware layer through automation, with a strong emphasis on reliability, observability, and scale. It suits someone who enjoys hands-on work close to BMCs and switches while also building the tooling that keeps production healthy.
Responsibilities
- Operate and scale Debian or Ubuntu-based bare-metal and virtualized Linux infrastructure across multiple sites. - Own the MAAS provisioning lifecycle, including region and rack controllers, PXE, commissioning, cloud-init, and node lifecycle automation. - Run production Kubernetes clusters end to end, covering upgrades, node pools, networking, storage, hardening, and incident troubleshooting. - Design and maintain multi-site networking spanning VLANs, L2/L3 routing, bonded interfaces, VPNs, firewalls, and DNS. - Automate provisioning and operations using Ansible, Bash or Python, and OpenTofu or Terraform in Git-based workflows. - Lead incident response, on-call coverage, SLI and SLO definition, and post-incident reliability improvements.
Requirements
- 5+ years of hands-on SRE, infrastructure, systems, or platform engineering experience. - Expert Linux administration on Debian or Ubuntu, with strong production MAAS and bare-metal provisioning skills. - Deep Kubernetes operations experience covering cluster lifecycle, networking, storage, and upgrades. - Solid network engineering across VLANs, L2/L3 routing, bonding, VPNs, firewalls, and DNS. - Strong automation background with Ansible plus Bash or Python, and familiarity with Terraform or OpenTofu. - Production observability experience with Prometheus and Grafana, plus incident response and on-call practice. - Comfort working autonomously in a fast-paced, engineering-driven environment.
Nice to have
- GPU infrastructure or GPU-heavy Kubernetes and bare-metal experience. - Proxmox VE with ZFS or Ceph and GPU passthrough. - VictoriaMetrics, VictoriaLogs, NetBox, Vault, SOPS, Atlantis, Cloudflare APIs, UniFi, service mesh, or Go tooling.
Benefits and work setup
- 100% remote with flexible hours, working in or near CET (±2h). - High-impact role with significant technical ownership and autonomy. - International engineering team focused on automation, reliability, and infrastructure at scale. - Opportunity to shape the architecture and operational foundations of a growing cloud platform.