Remote job
DevOps Engineer
Job details
About this role
Role overview A startup-focused infrastructure engineering team needs a DevOps Engineer to anchor the platform side of client engagements, from first commit through scaled production. The role combines hands-on engineering with consulting — you design the cluster, build the deploy path, set up observability, carry the pager, and hand over something the client's team can actually run without you. Work spans early-stage POC environments through mature multi-cloud setups, including air-gapped contexts.
Responsibilities - Author infrastructure-as-code with Terraform (plus Helm and GitOps via Argo CD or Flux) so every environment is reproducible from a blank account. - Operate Kubernetes clusters across managed services (EKS/GKE/AKS), on-prem, multi-cloud, and air-gapped deployments. - Build CI/CD pipelines using GitHub Actions, GitLab CI, or equivalents, incorporating signed images, SBOMs, and dependency scanning. - Implement observability stacks with OpenTelemetry, Prometheus, Grafana, and structured logs tied to meaningful SLOs. - Write internal tooling, operators, and automation in Go or Python. - Produce runbooks, design docs, and incident response processes that transfer capability to the client. - Support control mapping for SOC 2 and ISO 27001 audits.
Requirements - 3–5 years in DevOps, SRE, or infrastructure engineering with production systems you personally owned. - Real Kubernetes depth — scheduling, networking, RBAC, CRDs, and recovery behavior. - Terraform or equivalent IaC as a default reflex, not a fallback. - Strong in at least one of AWS, GCP, or Azure; functional in the other two across IAM, networking, releases, and cost. - Solid Linux and networking fundamentals for debugging below the abstraction. - Demonstrated pager experience for something that actually mattered. - Clear technical writing — runbooks and design docs are core deliverables. - Comfort working remotely with startup clients across time zones.
Nice to have - Go or Python for internal tooling and operators. - Multi-tenant SaaS platform work or direct compliance audit experience. - Open-source contributions in the cloud-native ecosystem. - GPU scheduling or inference-serving infrastructure exposure. - Early-stage startup background — founding team, early engineer, or your own shipped product.