Remote job
Data Engineer
Job details
About this role
Role overview This position leads end-to-end data modernization on Microsoft Azure, turning raw client requirements into scalable, governed data pipelines that fuel analytics and AI workloads. The work centers on Azure Synapse Dedicated SQL Pools and Fabric-style data hubs, with a strong emphasis on preparing data for downstream machine learning and natural-language search use cases. It suits an autonomous engineer who enjoys owning the full lifecycle from ingestion through optimization.
Responsibilities - Design, build, and tune ETL/ELT pipelines using advanced T-SQL and Python across Synapse Dedicated SQL Pools and Azure Data Hub environments. - Develop medallion-style layered schemas (Bronze, Silver, Gold) tailored for high-performance analytics and AI-ready consumption. - Embed AI capabilities into data operations, including automated incident triage, cost-saving recommendations, and proactive pipeline monitoring. - Implement data quality checks, observability frameworks, and governance controls across warehouses, lakes, and SQL/NoSQL stores. - Orchestrate workflows with tools such as Airflow, Prefect, or Dagster, and partner with AI/ML engineers to ground LLM applications in trusted data through vector embeddings and semantic metadata. - Optimize storage and compute spend across the Azure stack using FinOps automation and workload tuning, while troubleshooting issues independently.
Requirements - Deep expertise in T-SQL and hands-on experience with Azure Synapse Dedicated SQL Pools, Azure Data Hub, and relational/NoSQL databases. - Strong Python proficiency for data manipulation, pipeline development, and AI orchestration, including libraries like PySpark and Pandas. - Practical knowledge of the Azure data stack, including Synapse Pipelines and ADLS Gen2, at enterprise data-warehouse scale. - Familiarity with AI integration patterns such as prompt engineering, vector databases, and semantic layers for natural-language querying. - Solid grasp of data warehousing concepts, dimensional modeling, and data quality principles. - Ability to work independently, take ownership of infrastructure components, and uphold data security and compliance standards.