Remote job
Senior Platform Engineering Manager
Job details
About this role
Role overview A senior leadership role heading the Cloud Platform and Site Reliability Engineering functions for a human-data infrastructure company. The position blends people management with deep architectural influence over infrastructure, developer experience, and reliability practices. The manager is expected to set technical direction, embed SRE culture, and treat the internal platform as a product in its own right.
Responsibilities - Lead, coach, and develop one or more Platform teams, fostering a culture of ownership and high performance. - Drive the strategic roadmap for infrastructure and platform services, including managed Kubernetes clusters, with a focus on operational excellence and continuous improvement. - Embed SRE practices across the organization, owning service availability through SLOs, SLAs, error budgets, observability, and incident remediation. - Own the internal developer platform, including golden paths, self-service infrastructure, and tooling, and track adoption through platform-product metrics. - Ensure cloud environments remain secure and compliant with standards such as SOC 2 and ISO 27001. - Translate complex technical trade-offs into clear narratives for non-technical stakeholders.
Requirements - Proven track record leading infrastructure and SRE teams, scaling both platforms and people. - Several years of deep hands-on expertise with Google Cloud Platform, Kubernetes, and Infrastructure-as-Code tools such as Terraform, Terragrunt, or Crossplane. - Experience with GitOps workflows and tooling such as ArgoCD. - Strong grasp of distributed application architecture, observability principles, and incident management. - Demonstrated ability to balance platform stability and reliability with developer productivity and delivery velocity. - Strong written and verbal communication skills for working across technical and business audiences.
Nice to have - Familiarity with the broader stack including Python, TypeScript, Django, MongoDB, Postgres, Elasticsearch, DynamoDB, CircleCI, GitHub Actions, Celery, EventBridge, and Datadog.
Benefits and work setup - Remote working within a globally distributed engineering organization.