Remote job
Principal Software Engineer
Job details
About this role
Role overview
Drive a code-first transformation of site reliability for globally distributed, highly available systems. This principal-level individual-contributor role combines software engineering, architecture, incident leadership, cloud migration strategy, and technical mentorship to create durable platforms, self-healing services, and reliability practices that scale.
Responsibilities
- Lead and contribute code that makes reliability a product capability rather than a reactive operations function. - Design durable platforms and self-healing systems that reduce manual intervention and recurring incidents. - Act as a senior incident commander during high-severity global outages, bringing structure to stabilization and follow-up engineering. - Define resilient cloud-migration patterns and multi-active architectures designed to withstand regional failures. - Establish reliability standards such as meaningful error budgets and use them to inform product and engineering priorities. - Mentor senior and staff reliability engineers to approach infrastructure as software architecture. - Shape developer-experience strategies and internal platforms that provide resilience and infrastructure abstractions through standard workflows.
Requirements
- 15+ years in large-scale distributed systems, software engineering, or infrastructure, with demonstrated architectural leadership. - Deep proficiency in Go or Python and a track record of writing production-grade automation services. - Experience architecting global active-active distributed systems at very large scale while remaining hands-on as an engineering peer. - Production expertise with Kubernetes and Terraform across complex, multi-tenant cloud environments. - Experience serving as lead incident commander for critical global outages. - Experience defining developer-experience strategies or contributing to internal developer platforms.
Nice to have
- Experience leading on-premises-to-cloud migrations using Microsoft Azure, including Azure Kubernetes Service and Azure traffic routing. - Hands-on experience designing global distributed tracing with OpenTelemetry.
Benefits and work setup
- Remote role performed primarily from a designated remote work location. - Employment is unavailable in Alaska, Hawaii, Maine, Mississippi, North Dakota, South Dakota, Vermont, West Virginia, and Wyoming.