Remote job
Staff Software Engineer, Replication Foundations
Job details
About this role
Role overview Join a Replication Foundations team within a cloud global services organization as a Staff Software Engineer focused on a distributed systems replication stack that powers an open source workflow engine and its hosted cloud counterpart. The work centers on correctness-critical infrastructure that enables high-availability namespaces, cross-cluster and cross-region failover, and workload migration between self-hosted and cloud deployments. This is a senior, cross-functional technical leadership role spanning architecture, design, implementation, rollout, and operational stewardship.
Responsibilities - Set the technical direction for the open source replication stack, from problem framing through rollout and operational support. - Lead design and implementation of replication protocols supporting high-availability namespaces, cross-cluster and cross-region replication, and migrations between clusters in either direction. - Drive scalability and reliability work such as multi-cell namespaces, spanning namespaces across multiple clusters, load distribution, and hot-spot mitigation. - Define and communicate system-level guarantees covering consistency models, ordering, idempotency, failure recovery, performance, and operational behavior. - Identify architectural risks and shape the technical roadmap for replication capabilities that underpin current and future cloud products. - Partner across cloud enablement, product, and engineering teams to align replication foundations with customer and product needs, including leading design reviews.
Requirements - Significant experience building production distributed systems where correctness, consistency, and failure handling are first-order concerns. - Deep familiarity with replication protocols, consensus, data consistency, and ordering guarantees. - Demonstrated ability to lead complex initiatives end to end, from architecture through operations. - Strong communication skills, with the ability to explain complex designs and trade-offs to both technical and cross-functional audiences. - Track record of influencing technical direction across teams, building alignment without direct authority, and mentoring engineers. - A thoughtful, curious approach to understanding how systems behave under load, failure, and shifting workload conditions.
Nice to have - Hands-on work designing or maintaining replication protocols or data-plane infrastructure. - Experience with multi-cluster or multi-region architectures, including active-active or active-passive systems. - Familiarity with database internals, log-based replication, or event-sourced systems. - Prior contributions to large open source projects or distributed systems infrastructure.