Remote job
Staff Platform Engineer (Kubernetes)
Job details
About this role
Role overview A staff-level platform role owning the control plane for a multi-cluster Kubernetes fabric that orchestrates GPU workloads across regions and providers. The work covers fleet lifecycle, scheduling, isolation and graceful drain when capacity runs short, and multi-tenancy is treated as a first-class concern because customers share expensive hardware. This is a position with genuine architectural latitude, expected to set direction rather than only implement someone else's design.
Responsibilities - Own the multi-cluster control plane: fleet lifecycle, upgrade choreography, drift detection and configuration consistency - Design and build custom resources and operators covering GPU allocation, sharing, reclamation and eviction - Enforce multi-tenant isolation across compute, network and storage so noisy or hostile neighbours stay contained - Own scheduling behaviour for GPU workloads, including queueing, prioritisation, preemption and graceful drain - Drive the GitOps and change-management model for infrastructure so changes are reviewable, auditable and reversible - Build cost-allocation and utilisation signals that show where capacity actually goes - Set technical direction for the platform group and mentor engineers through design review rather than by decree
Requirements - Several years operating Kubernetes in production at scale, including having been responsible when it went wrong - Experience authoring controllers or operators, with a working understanding of the reconciliation loop - Willingness to work primarily in Go - Real multi-tenancy experience covering namespaces, network policy, resource governance and the isolation boundaries that matter - Fluency with infrastructure as code and GitOps-style delivery - Sound judgement about operational blast radius, staged rollout and rollback - Ability to write a design document that a sceptical reviewer can disagree with productively
Nice to have - Direct experience scheduling GPU or other accelerator workloads in Kubernetes - Familiarity with device plugins, GPU operators or topology-aware scheduling - Background in multi-cloud or hybrid estates spanning owned hardware and public cloud - Experience with service mesh, advanced networking or eBPF-based tooling
Benefits and work setup - Remote-eligible, full-time platform role with competitive compensation