Remote job
Staff Software Engineer, Compute and Network
Job details
About this role
Role overview Set the technical direction for the compute and networking platform that underpins a microservices architecture powering AI-driven personalized marketing. This hands-on staff role focuses on building reusable cloud components, clear interfaces, and self-service tooling that make infrastructure safer and easier to consume.
Responsibilities - Shape the technical direction and architecture of the platform alongside engineering leadership, prioritizing investments. - Partner with engineering teams to understand workflows and pain points, then deliver platform software, APIs, and self-service tools that reduce friction. - Lead complex, cross-team initiatives, turning ambiguous requirements into clear designs, delivery plans, and measurable outcomes. - Evolve Kubernetes and networking infrastructure across multiple cloud accounts and networks to support growth, resilience, and efficient resource use. - Own and improve Infrastructure-as-Code orchestration and reusable cloud components with consistent interfaces and guardrails for cloud services. - Diagnose complex production issues and drive lasting improvements in system design, observability, and operational practice. - Improve platform reliability, scalability, and cost efficiency while reducing operational overhead. - Raise the engineering bar through hands-on development, design and code reviews, mentorship, and technical guidance. - Identify and establish effective AI-assisted engineering practices that boost productivity without compromising quality, security, or reliability.
Requirements - Significant experience designing, building, and operating large-scale infrastructure or distributed systems, with demonstrated technical leadership across teams. - Strong software engineering background, including building maintainable software, tooling, or APIs that solve infrastructure problems. - Demonstrated ability to clarify ambiguous problems and lead complex technical initiatives through delivery and production ownership. - Customer empathy for internal developer users, with the ability to translate feedback into platform improvements and assess success through adoption and outcomes. - Substantial production Kubernetes experience, including architecture, networking, and failure modes. - Strong understanding of networking fundamentals across L3 through L7, diagnosing complex connectivity, routing, DNS, load-balancing, and service-to-service communication issues. - Experience in release engineering, including CI/CD pipeline design, automated delivery, staged rollouts, and reliable rollback and recovery. - Experience using observability tooling to understand system behavior and guide reliability improvements. - Comfort applying AI-assisted development tools to accelerate implementation, testing, debugging, and documentation, with sound judgment on generated code and sensitive data. - Commitment to mentorship, constructive feedback, and continuous learning.
Nice to have - Background operating production networking infrastructure on AWS with Kubernetes, Istio, Argo CD, Terraform, and Cloudflare at the edge.
Benefits and work setup - Competitive perks and benefits package including health and wellness offerings and equity. - Distributed global workforce with employee hubs in major cities and remote-friendly practices.