Remote job
Senior Data Engineer
Job details
About this role
Role overview A senior, hands-on data engineering role focused on building and operating the analytics backbone that powers internal insights and customer-facing products. The position owns a CDC-fed medallion data lake, a Postgres warehouse, and the integration of a graph database, while extending production generative AI workflows.
Responsibilities - Own the production data lake end-to-end, including CDC ingestion, medallion layers on Apache Iceberg, the Trino query layer, and the table catalog. - Operate and evolve the Postgres data warehouse, including schema design, performance tuning, and access controls. - Integrate the graph database with the warehouse, lake, and primary databases, migrating legacy pipelines onto a single governed integration layer. - Maintain and evolve Dagster-orchestrated dbt pipelines with sensor-triggered builds, data quality tests, and branch-based versioning. - Operate the BI and reporting layer, enforcing per-user policy-based data access at the query engine rather than only at the dashboard. - Support and extend production genAI workflows such as embeddings, similarity search, and LLM-based extraction and classification.
Requirements - 5+ years in data engineering, analytics engineering, or data platform engineering. - Advanced SQL and relational database experience with Postgres and MongoDB, plus hands-on production graph database experience. - Experience with open table formats and medallion architectures such as Apache Iceberg, distributed SQL engines like Trino or Presto, and streaming or CDC pipelines (e.g., Kafka, Debezium, Flink). - Strong Python skills for pipelines and automation, plus experience using dbt orchestrated by a modern scheduler like Dagster or Airflow. - AWS infrastructure experience via infrastructure-as-code such as Terraform, along with CI/CD, GitOps, and Docker. - Background in analytics data modeling, metric definitions, automated monitoring, and data security practices.
Nice to have - Familiarity with BI tooling supporting per-user policy-based access, policy engines like OPA, and identity platforms such as Keycloak. - Experience with multi-region data residency patterns, including separate EU and China data handling.
Benefits and work setup - US base salary of $135,000–$165,000 with a 10% annual bonus, incentive stock options, and work-from-home stipends. - Medical, dental, and vision coverage with 90% of employee premiums and 60% of spouse or dependent premiums covered, 401(k) with up to 4% match, fully paid parental leave, unlimited PTO, and 12 paid company holidays.