Remote job
Cloud / Platform Engineer (Kubernetes, GCP, AWS)
Job details
About this role
Role overview
Lead the infrastructure behind an on-demand application deployment platform that turns generated apps into live URLs, running either on shared cloud infrastructure or inside customer-owned AWS and GCP environments. The mandate is to take a working but shell-heavy production stack and convert it into something automated, reliable, and operationally calm.
Responsibilities
- Operate the Kubernetes foundation: a highly available control plane, an internal container registry, deployment topology decisions, and the workload manifests everything ships from. - Engineer and maintain deployment pipelines that target both GCP (Cloud Run, Cloud Build, GCS) and AWS (CodeBuild, ECR, App Runner/ECS), including replacing the lingering customer-issued IAM access key pattern with a keyless, fail-closed impersonation model. - Separate short-lived preview environments from long-running production workloads and reconcile the multiple sources of truth that currently determine app lifecycle state. - Own the build-to-release path: image building, promotion between environments, controlled rollout, fast rollback, and clear in-pod version visibility. - Stand up logs and metrics that humans can navigate, alerts tied to real signals, and cloud-spend controls that grow with customers instead of incidents. - Harden secrets handling, network policy, TLS termination, served-app security headers, and the tenant isolation layer at the infrastructure boundary. - Co-build a real on-call rotation with documented escalation and a roster that does not lean on one person.
Requirements
- Four or more years running production cloud infrastructure, platform engineering, or SRE work. - Deep, incident-tested Kubernetes experience, not just coursework familiarity. - Working expertise in AWS and/or GCP, with confidence navigating IAM in either cloud. - Infrastructure-as-code practice plus practical shell and Python skills. - Demonstrated CI/CD ownership and a preference for low-drama deployments.
Nice to have
- Multi-tenant infrastructure background. - Cost optimization or FinOps experience. - Terraform fluency. - Datadog or equivalent observability stack. - Prior work hosting customer workloads, especially with on-prem or hybrid delivery.