Remote job
Lead Data Engineer - Identity
Job details
About this role
Role overview
Lead the design and operation of identity data systems supporting cross-device, cross-surface targeting, measurement, and attribution. The role combines hands-on data engineering with technical leadership: you will shape a new multi-source identity graph, make partner onboarding repeatable, and prepare audience data for self-service discovery. You will also establish engineering standards and help a team deliver reliable, privacy-aware systems.
Responsibilities
- Define the architecture and rollout plan for identity synchronization, identifier translation, clustering, and opt-out handling. - Build reusable ingestion patterns for partner and client feeds across web, connected TV, mobile, and cleanroom environments. - Develop the data layer for audience creation, activation, lifecycle state, and reporting. - Improve observability by standardizing freshness and quality signals, alerts, incident context, and operational runbooks. - Set expectations for testing, cost efficiency, reusable patterns, and technical decision records through code reviews and knowledge sharing. - Work with Product, Data Partnerships, and Data Science to turn ambiguous requirements into a sequenced roadmap.
Requirements
- Experience designing and owning large, interdependent data systems, including at least one platform built from the ground up. - Experience leading engineers while remaining hands-on with technical direction, implementation, reviews, and development. - Strong command of Python, Airflow, Spark, and disciplined SQL development for Snowflake, with attention to performance and cost. - Practical experience with AWS, Kubernetes, infrastructure diagnostics, and third-party APIs in ingestion pipelines. - Ability to use AI-assisted engineering tools effectively and structure codebases so they remain understandable to both people and automated tools.
Nice to have
- Identity resolution, device or household graph, or other matching experience in advertising technology. - Knowledge of privacy and consent requirements, including opt-outs, deletion workflows, GDPR, and CCPA. - Experience with data cleanrooms, GitHub Actions, ArgoCD, VictoriaMetrics, Prometheus, or Grafana. - Production experience with Iceberg, Kafka or comparable streaming systems, Aerospike or another low-latency store, or ClickHouse.