Remote job
Senior Data Engineer
Job details
About this role
Role overview A senior data engineering role focused on maturing a Databricks-based lakehouse platform that ingests, transforms, and serves large volumes of engineering data. The work spans ingestion, transformation, storage, governance, and serving layers, with a goal of turning messy pipelines into durable, well-modeled platform architecture. It is suited to someone who wants to define how a modern lakehouse actually operates end to end.
Responsibilities - Build and maintain data pipelines and datasets in Databricks and Delta Lake, improving reliability, performance, and operational visibility. - Establish and enforce clear Bronze, Silver, and Gold layer responsibilities, including standards for schema evolution, transformation ownership, retention, and promotion between layers. - Design durable canonical models for core entities and relationships, partnering with application and analytics teams to ensure consistent definitions downstream. - Build and improve batch and incremental pipelines using Databricks, Airflow, Spark, and cloud object storage, prioritizing idempotency, scalability, observability, and recoverability. - Work with catalog and governance tooling to establish lineage, ownership, schema standards, quality checks, and discoverability. - Create reliable patterns for serving curated data into systems like ClickHouse and other future destinations without tight coupling to any single database.
Requirements - Extensive hands-on experience with Databricks, Spark, Delta Lake, or a comparable lakehouse platform, beyond writing notebooks. - Strong data engineering fundamentals including partitioning, incremental processing, schema evolution, distributed execution, file formats, and the performance characteristics of large analytical datasets. - Solid data modeling skills covering canonical entities, relationships, grain, dimensional modeling, and the boundary between platform and consumer models. - A pipeline reliability mindset, with pipelines designed to be observable, retryable, idempotent, and easy to debug on failure. - Cloud fluency across object storage, compute, networking, IAM, and managed data services in a modern cloud architecture. - A pragmatic, platform-builder approach that balances standards with shipping practical solutions and iterating.
Nice to have - Prior experience building or migrating to a medallion-style lakehouse architecture. - Familiarity with Databricks Unity Catalog, OpenMetadata, or other governance and lineage platforms. - Experience implementing CDC pipelines from PostgreSQL, RDS, or Aurora. - Hands-on work with Airflow or another production workflow orchestration platform. - Experience moving analytical data into serving systems such as ClickHouse, Snowflake, or BigQuery. - Helping introduce data contracts, canonical schemas, or platform-wide data quality standards.
Benefits and work setup - Occasional travel may be required. - Applicants must be authorized to work for any employer in the US; visa sponsorship is not available at this time.