Remote job
Platform Infrastructure Engineer (SRE Core)
Job details
About this role
Role overview
This is a Platform Infrastructure Engineer role on the Site Reliability Engineering (SRE) team that builds and operates the core infrastructure services for a cloud-native platform. The position spans multiple regions and environments across major cloud providers, with a strong emphasis on infrastructure as code, multi-region resilience, comprehensive observability, and security-first design. The team is globally distributed, uses AI-assisted development tooling in its standard workflow, and operates with a 24x7 on-call rotation.
Responsibilities
- Implement, deploy, and maintain VM and Kubernetes infrastructure across multiple cloud providers and regions, spanning dozens of clusters across development, staging, and production environments - Build and maintain infrastructure as code using Terraform modules and orchestration tooling such as Spacelift, provisioning networking, compute, storage, and security components - Implement and maintain observability using Grafana Cloud, Prometheus/Mimir, and OTel collectors, designing dashboards and alerting rules across all platform components - Manage certificate lifecycle, DNS automation, ingress controllers, and service mesh networking - Partner across Engineering, Product, Compliance, and Security teams on capacity planning, disaster recovery, and architectural decisions - Identify and eliminate operational toil through automation, scripts, CI/CD pipelines, and AI-assisted coding tools - Participate in a 24x7 on-call rotation as part of a globally distributed team, responding to incidents and driving post-incident reviews
Requirements
- Bachelor's degree in Computer Science or a related technical field, or equivalent practical experience - Proficiency in programming and scripting languages, particularly Python, Bash, and Go - Understanding of network topologies, communication protocols (TCP/IP, HTTP/S, UDP, TLS), and enterprise connectivity solutions - Production Kubernetes expertise including cluster administration, RBAC, networking, workload management, and troubleshooting - Proven experience with Terraform for infrastructure provisioning and management - Knowledge of major cloud platform services such as GKE, VPC networking, Cloud DNS, Artifact Registry, Secret Manager, and IAM
Nice to have
- Familiarity with AI-assisted development and code-review tooling, including LLM-based assistants used as part of standard engineering workflows