Remote job
Senior Site Reliability Engineer - Volcano
Job details
About this role
Role overview
A greenfield internal developer platform initiative is seeking a Senior Site Reliability Engineer to act as the reliability lead for a brand-new platform spanning preview environments, edge deployments, managed PostgreSQL, authentication, realtime APIs, and storage. Sponsored by the Office of the CTO, the role partners directly with engineering leadership to shape the platform's reliability posture from day one. It is a high-visibility position with significant influence on architectural decisions across a multi-region, multi-tenant stack.
Responsibilities
- Define and own SLOs, error budgets, and incident response practices for every service on the platform, including the control plane, edge deployments, managed databases, auth, and storage layers. - Design and build the underlying Kubernetes infrastructure across multiple regions, covering networking, the data plane, and the edge deployment pipeline that powers backend-as-a-service capabilities. - Establish GitOps and CI/CD patterns using ArgoCD, Helm, and Terraform or Terragrunt, including canary release pipelines and on-demand preview environment provisioning that other teams will adopt. - Operate and harden multi-tenant PostgreSQL clusters, Redis caching, and object storage, with a focus on tenant isolation, performance, and disaster recovery. - Build observability in from day one using Datadog, Prometheus, and Grafana, including SLIs, dashboards, alerts, and runbooks created before services go live. - Collaborate with product engineering, security, and the CTO office to embed reliability and compliance into the architecture rather than bolting them on later, and evaluate emerging technologies such as edge runtimes, serverless compute, and AI-native infrastructure.
Requirements
- Bachelor's degree in Computer Science or equivalent experience, with a track record at Staff or Principal IC level in SRE or platform engineering. - Demonstrated experience building SRE or platform practices for developer-facing platforms, PaaS, or SaaS products, ideally from a greenfield stage. - Deep Kubernetes expertise across multi-tenant cluster design, networking (CNI, service mesh, ingress), autoscaling, and security hardening. - Hands-on experience with managed data services, including PostgreSQL, Redis, and object storage at scale. - Proficiency with GitOps tooling (ArgoCD, Helm, Terraform or Terragrunt) and the observability stack mentioned above.
Nice to have
- Experience evaluating or adopting emerging infrastructure categories such as edge runtimes, serverless compute, and vector databases. - Comfort working in a strategic, cross-functional capacity alongside engineering leadership and security teams.