Remote job
Senior/Staff Platform Engineer
Job details
About this role
Role overview
A hands-on senior or staff platform engineering role focused on building, operating, and evolving large-scale production infrastructure. The position spans Kubernetes, Linux, cloud, networking, observability, CI/CD, and reliability, with production coding in Go, Python, or Java to support platform and tooling needs. It is a highly autonomous, remote role aligned to Pacific Time hours, with on-call rotation roughly every four to five weeks, and is open to candidates based in Brazil, Mexico, or Canada.
Responsibilities
- Design, build, operate, and improve production Kubernetes platforms, owning cluster architecture, networking, workload isolation, security, upgrades, scaling, and reliability. - Troubleshoot Kubernetes below the application layer, including CNI/networking, scheduling, node behavior, resource constraints, controllers, and cluster-level failures. - Operate large, highly available infrastructure across cloud, hybrid, virtualized, and bare-metal environments, and optimize it for reliability, performance, scalability, and efficiency. - Write and maintain production tooling and automation in Go, Python, or Java, and build internal services, APIs, and operational tooling. - Improve incident detection, time to diagnosis, and recovery through better observability and runbooks. - Drive complex infrastructure initiatives from problem definition through design, implementation, and production operation, and contribute to architecture discussions, RFCs, and design reviews. - Mentor engineers and raise the technical bar across the team.
Requirements
- 10+ years in platform, SRE, infrastructure, or DevOps engineering, with significant hands-on production operations experience. - Demonstrated experience building and operating production Kubernetes platforms, not just deploying apps onto existing clusters. - Production programming experience in Go, Python, or Java. - Strong Linux fundamentals and production systems troubleshooting skills. - Deep understanding of Kubernetes internals, including networking/CNI, NetworkPolicy, scheduling, RBAC, and cluster behavior. - Track record of independently owning ambiguous technical initiatives and communicating complex tradeoffs to stakeholders. - Degree in Computer Science or related field, or equivalent practical experience.
Nice to have
- Cloud migration experience. - Background operating across regulated or hybrid enterprise estates. - Prior mentoring or informal technical leadership at staff level.
Benefits and work setup
- Remote role with Pacific Time working hours and on-call rotation approximately every 4–5 weeks. - High autonomy with direct technical stakeholder access. - Open to candidates in Brazil, Mexico, or Canada.