Remote job
Senior DevOps Engineer
Job details
About this role
Role overview
A senior infrastructure engineer is needed to take full ownership of a self-operated platform built on Proxmox and Ceph, spanning bare metal and multiple cloud providers. The platform carries very high transaction volumes, and the new hire will own it end to end, drive FinOps savings, and split development workloads into a separate cluster without disrupting production.
Responsibilities
- Own the Proxmox and Ceph platform across capacity planning, upgrades, performance tuning, and incident response. - Operate infrastructure across OVH, Hetzner, and bare metal hardware, alongside workloads running in AWS and GCP. - Lead resource rightsizing, identify over-provisioned services, and demonstrate savings in the bills. - Drive the separation of development workloads into a dedicated cluster with zero downtime for production. - Keep all infrastructure described as code using Terraform, Ansible, ArgoCD, and GitLab CI, with no manual production changes. - Own reliability practice: on-call rotation, incident response, root cause analysis, and runbook maintenance. - Support the underlying data layer powering the product, including PostgreSQL, ClickHouse, and Aerospike.
Requirements
- 10 or more years working hands-on with infrastructure and physical hardware, beyond managed-only cloud experience. - 5 or more years running Kubernetes and cloud platforms in production. - Direct experience with bare metal and providers such as OVH and Hetzner, plus AWS or GCP. - Strong proficiency with Terraform, Ansible, ArgoCD, and GitLab CI as the default way of describing infrastructure. - Practical FinOps experience with concrete before-and-after cost numbers for changes you have made. - A root-cause mindset that follows problems to a fix, including work outside your formal scope. - Direct communication style that surfaces risks and bad news early rather than after an incident.
Nice to have
- Production experience specifically with Proxmox and Ceph. - Operating PostgreSQL, ClickHouse, or Aerospike at scale. - Splitting shared environments into isolated clusters. - Building observability and on-call practice from scratch. - Background in high-traffic B2B SaaS.
Benefits and work setup
- Compensation tied directly to results. - Fully remote flexibility, with the option to work onsite in Belgrade or London. - Professional support covering language learning, sports, health insurance, and laptop coverage.