Job details
About this role
Role overview A senior engineering role focused on building, extending, and operating a near real-time streaming data platform for operational data across the business. The position is hands-on and end-to-end: the same engineer who designs the platform also owns it in production, including latency and throughput SLOs, replay/backfill, dead-letter triage, and connector failure recovery. The work sits in a regulated environment where governed access to protected health information (PHI) is a first-class concern.
Responsibilities - Run the streaming platform in production, owning latency and throughput SLOs, monitoring and alerting, replays, per-source backfills, connector-failure recovery, dead-letter triage, and tuning under spiky, batch-driven load. - Build and extend streaming pipelines that ingest CDC events from operational databases and SaaS sources into a canonical, contract-validated form. - Transform and enrich data in Spark (Structured Streaming or dbt-on-Spark micro-batch), including cross-stream joins that correlate events into unified lifecycle entities. - Enable secure, governed data access over PHI through classification, row/column controls, and access policies applied as data is served to consumers. - Make pipelines reliable and observable through idempotent, replayable design, data-quality validation, runbooks, and observability the broader data team can rely on. - Provision the platform as code (streaming, processing, storage, and catalog services on AWS) using CDK with CI/CD for data pipelines, right-sized for cost against latency SLOs. - Partner with upstream producers on source changes and contracts, downstream consumers on access and data needs, and stakeholders on translating requirements into platform capabilities.
Requirements - 6+ years building and operating production data systems, with demonstrated ownership of streaming or event-driven pipelines including on-call, incident response, SLOs, runbooks, and recovery. - Deep, production experience with Apache Kafka, including partitioning, consumer groups, consumer-lag and broker-health troubleshooting, exactly-once or idempotent semantics, schema registry, and replay/backfill under load. - Strong hands-on Apache Spark experience (PySpark) for streaming and batch transformation in production. - AWS-native data engineering across streaming, processing, storage, and catalog services, with infrastructure-as-code (AWS CDK preferred or Terraform), CI/CD for data pipelines, and cost awareness. - Comfort debugging distributed data pipelines (consumer lag, data skew, backpressure, late or out-of-order events) using observability tooling. - Expert-level SQL and Python, experience building and consuming REST APIs, and a degree in Computer Science, Information Systems, or another quantitative field. - Hands-on use of AI-assisted and agentic development tools, with an opinion on where AI adds value, plus familiarity with CDC tools (e.g. Debezium), open table formats like Iceberg or Delta Lake, data contracts and schema governance, and health-insurance domain knowledge including HIPAA PHI.
Nice to have - Experience with Apache Flink or other stateful stream processors. - Knowledge of JVM-based languages such as Kotlin or Java. - Familiarity with serving data to AI and agentic consumers as low-latency context or inputs. - Previous venture-backed startup experience.
Benefits and work setup - Alternative medicine coverage, flexible PTO, up to 16 weeks of paid parental leave, paid holidays, a 401k program, transportation perks, and education reimbursement. - Two days of paid paw-ternity leave, on top of standard health and wellness benefits. - Mac or Linux command-line environment.