Remote job
Principal Data Engineer
Job details
About this role
Role overview
This is the top individual contributor role on a data platform team at a large direct-to-consumer healthcare company. The platform powers analytics, machine learning, and product engineering across telehealth, prescription, and wellness products used by millions of customers, so architectural decisions made here shape how the broader engineering organization works with data for years to come. The seat exists because capable Staff engineers already ship strong domain work, but no one currently owns whether those domains add up to a coherent, forward-looking platform.
Responsibilities
- Define the long-term platform architecture across ingestion, streaming, orchestration, transformation, and serving, choosing what to build, buy, or sunset, and in what sequence - Make the streaming decision: whether a Kafka-to-Flink-to-warehouse pipeline stays scoped to analytics or graduates into production and model-evaluation infrastructure, with all the operational and financial consequences that follow - Set the data contracts, service boundaries, and interfaces that govern how adjacent teams consume from and publish to the platform - Define what "trustworthy data" means in practice, including quality and freshness guarantees, dataset tiering, and the accountability model when guarantees slip - Shape the platform's economics, from storage and compute strategy to reservation design, and create the framework that keeps growing data consumption honest - Design a self-service model so analytics and ML teams can move without the data team becoming a bottleneck, paired with guardrails that protect production - Establish engineering standards (testing, CI/CD, observability, schema governance, infrastructure as code) that other teams adopt because they are good, not because they are mandated - Chair Architecture Review and serve as the decision-maker of record for cross-team, multi-system, and cost-impacting changes, including systems owned by other organizations
Requirements
- Production experience operating Databricks, Unity Catalog, and Delta Lake at meaningful scale - Hands-on CDC patterns and Flink for real-time processing, especially where streaming serves both analytics and production or model-serving paths - PySpark and SparkSQL for large-scale batch and streaming workloads - Track record leading a BigQuery-to-Databricks Lakehouse migration or an equivalent warehouse migration end to end - MLOps collaboration on model training pipelines, feature stores, experimentation infrastructure, or data foundations behind AI-driven product features - Go or Python experience building Kafka producers and consumers - Background in direct-to-consumer healthcare or telehealth with HIPAA and GDPR obligations - Applied SOX compliance controls in a data engineering context
Benefits and work setup
- Competitive salary and equity compensation for full-time roles - Unlimited PTO, company holidays, and quarterly mental health days - Comprehensive medical, dental, and vision coverage plus parental leave - Employee Stock Purchase Program - 401(k) with employer match - Offsite team retreats