Remote job
Middle Data Engineer
Job details
About this role
Role overview Join a data engineering team supporting a large-scale platform that connects insurers, repair facilities, parts suppliers, and connected-car service providers across the vehicle ecosystem. The position centers on designing ETL/ELT pipelines, maintaining structured data lakes, and curating feature stores that feed downstream machine learning models for subrogation. It suits a mid-level engineer who combines hands-on technical work with strong cross-team collaboration.
Responsibilities - Design, build, and tune ETL/ELT pipelines that transform customer and internal data into well-modeled analytical datasets - Partner with markets and product teams to deliver reporting tools and dashboards that inform business strategy - Develop and maintain data lakes and feature stores consumed by ML engineers building subrogation models - Apply medallion architecture patterns (bronze, silver, gold) with schema management and incremental processing - Implement data quality frameworks and debug issues that span distributed systems - Use infrastructure-as-code tooling to provision and manage cloud data services
Requirements - Strong proficiency in Python and SQL - Hands-on experience with Spark or comparable distributed processing frameworks - Practical experience with AWS data services such as Glue, S3, Step Functions, Athena, EMR, DynamoDB, and SQS - Familiarity with streaming platforms like Kafka or Spark Streaming - Working knowledge of MongoDB and other document-oriented data sources - Comfort with infrastructure-as-code using CloudFormation or Terraform
Nice to have - Background in subrogation, claims processing, insurance, or financial technology
Benefits and work setup - In-office, hybrid, or remote flexibility depending on location - Medical healthcare coverage - Recognition program and ongoing learning reimbursement - Well-being program and team events - Sports compensation and referral bonuses - Top-tier hardware provided