Remote job
Staff Platform Engineer
Job details
About this role
Role overview
An internal engineering platform team is hiring a Staff Platform Engineer to lead the software and infrastructure that enables product engineering teams to build, deploy, and operate services at scale. The role combines deep hands-on software engineering with cross-team technical leadership, focused on turning infrastructure, delivery, and reliability work into reliable, self-service platform products. The mission is to make engineering faster, safer, and more scalable by replacing manual operations with durable software and clear defaults.
Responsibilities
- Set technical direction and contribute to a long-term roadmap across cloud, networking, compute, storage, software delivery, and reliability engineering. - Lead architecture reviews and technical engagements across the software development lifecycle, from early design through production operation. - Design, build, and maintain platform services, APIs, CLIs, controllers, workflows, and automation in languages such as Go, TypeScript, Python, or similar, owning them through design, implementation, testing, deployment, and ongoing production support. - Build self-service capabilities that let engineering teams provision environments, deploy services, manage infrastructure, and respond to operational conditions without manual platform intervention. - Improve CI/CD and automated software delivery systems while increasing deployment safety and developer velocity. - Measure and reduce operational toil, coordinate response to significant incidents, run effective postmortems, and turn recurring issues into permanent reliability improvements. - Influence engineering standards across teams through design reviews, technical writing, prototypes, internal education, and mentorship.
Requirements
- 8+ years of professional experience in software engineering, infrastructure engineering, site reliability engineering, or an equivalent combination, including substantial ownership of production systems. - Strong software engineering ability in at least one language such as Go, Python, TypeScript, or Rust, with experience building maintainable services, CLIs, APIs, or developer tools. - Hands-on experience operating distributed systems in production, including failure diagnosis, capacity management, performance improvement, and high-availability design. - Practical experience with a major cloud provider (AWS, GCP, or Azure) and infrastructure-as-code tooling such as Terraform or an equivalent system. - Experience with containers and orchestration such as Kubernetes, EKS, GKE, AKS, Docker, or Nomad, including workload deployment, networking, security, and operational troubleshooting. - Experience owning cloud network security controls, including firewalls, DNS filtering, and network segmentation. - Track record of setting technical direction, leading architecture across multiple teams, and influencing decisions without relying on direct authority.
Nice to have
- Build tooling experience such as Bazel or Buck2. - Background in capacity planning, failure recovery, and the operational tradeoffs of running critical systems at scale. - Experience building secure-by-default systems and collaborating on identity, access control, secrets, supply-chain security, and policy enforcement. - Ability to learn unfamiliar domains quickly, including the infrastructure characteristics of data-intensive, event-driven, or AI/ML workloads.
Benefits and work setup
- Base salary range of $168,000 to $240,000 for New York, with a discretionary annual bonus and a new-hire equity grant. - Comprehensive health plans, 401(k) with company match, paid parental leave, and flexible time off. - Hybrid work model at hub office locations, with a remote workforce option for employees outside hub cities and required in-person onboarding.