Remote job
Principal Backend Engineer
Job details
About this role
Role overview A principal-level engineering position focused on scaling a high-throughput, low-latency real-time API to enterprise production scale. The work centers on designing resilient, fault-tolerant distributed systems that ship into external customer environments where the network, identity provider, observability stack, and operator are outside the team's direct control.
Responsibilities - Own the architecture and implementation of the real-time API built on Python/FastAPI and deployed to Kubernetes. - Engineer for resilience, designing graceful degradation under upstream failures, rate limits, partial outages, and retry storms while eliminating silent failure modes. - Tune Kubernetes autoscaling (HPA, KEDA, cluster autoscaler) for bursty production workloads across diverse deployment environments. - Build streaming pipelines that synchronize data between service, transactional, and analytical layers. - Architect services for portability across customer-managed identity, restricted network egress, and varied observability stacks without serverless assumptions. - Profile, benchmark, and optimize latency and throughput across the full stack, from request handling through connection pooling, caching, and downstream calls. - Set technical direction on distributed systems patterns, deployment architecture, and operational tooling.
Requirements - 10+ years of backend engineering experience with a principal or staff track record on high-performance distributed systems. - Deep Python and FastAPI expertise, including async patterns, performance profiling, and production-grade service development. - Proven experience building high-throughput, low-latency systems with strict latency requirements. - Strong Kubernetes experience, especially autoscaling, resource tuning, and production operations. - Resilience engineering skills: circuit breakers, retries with backoff and jitter, bulkheads, timeouts, idempotency, graceful degradation. - Hands-on streaming systems experience with Kafka, Pulsar, Kinesis, or similar platforms. - Strong observability instincts across structured logging, metrics, and tracing. - Clear technical communication and ability to collaborate across distributed engineering teams.
Nice to have - Databricks or lakehouse architecture experience. - Spark Structured Streaming familiarity. - Deploying into enterprise customer environments with strict security or compliance constraints, including air-gapped setups. - Regulated industry experience in healthcare or financial services. - Open-source contributions to distributed systems or streaming projects.