Remote job
Data Engineer
Job details
About this role
Role overview
Join a Databricks-native engineering team building the data foundation for a modern Master Data Management (MDM) platform that helps global enterprises unify and govern trusted data entities across customer, product, supplier, and asset domains. As a Data Engineer, you will design and optimize scalable, high-performance pipelines that power entity resolution, data quality, and real-time analytics. This is a self-directed, hands-on role suited to someone who thrives in a fast-paced environment where reliable, large-scale data systems are central to the product.
Responsibilities
- Design, build, and optimize robust batch and streaming data pipelines on the Databricks Lakehouse Platform, using Delta Live Tables and Change Data Capture patterns. - Translate business and product data requirements into efficient data models that support intelligent features, analytics, and downstream AI/ML use cases. - Ensure data reliability, consistency, and performance through rigorous pipeline design, testing, and Spark workload tuning. - Collaborate closely with AI/ML Engineers, Product Managers, and Data Scientists to deliver pipelines that enable analytics-ready and AI-ready data products. - Contribute to engineering best practices across monitoring, observability, version control, and automated testing for data workflows. - Work with MDM data structures and support integrations with modern cloud data services to deliver governed golden records.
Requirements
- Hands-on experience designing and operating scalable data pipelines, ideally on Databricks or equivalent Lakehouse platforms. - Strong proficiency with Apache Spark, Delta Lake, and modern data modeling patterns for batch and streaming workloads. - Solid understanding of Master Data Management concepts and the underlying data structures used to model core business entities. - Experience deploying data engineering workloads on major cloud platforms such as AWS or Azure. - Familiarity with CI/CD practices for data pipelines and infrastructure-as-code tools such as Terraform.
Nice to have
- Knowledge of MLOps practices and integrating data pipelines with machine learning workflows. - Experience with monitoring, observability, and automated testing frameworks for production data systems.
Benefits and work setup
- Full-time, remote position based in India. - Two openings currently available on the data engineering team.