Remote job
Sr. Data Integration Engineer
Job details
About this role
Role overview A senior engineering role focused on building and operating resilient batch and event-driven data pipelines on AWS for public-sector programs. The work blends hands-on Python development, cloud data engineering, infrastructure automation, and cross-team collaboration to turn raw data into trusted, production-ready datasets.
Responsibilities - Design and develop batch and event-driven pipelines using Python, AWS Glue, and supporting AWS managed services - Build and orchestrate Glue jobs, crawlers, workflows, triggers, and Data Catalog integrations for data discovery, transformation, governance, and publishing - Ingest and process structured and semi-structured data (including JSON and NDJSON) from files, APIs, databases, and streaming sources - Implement schema validation, normalization, enrichment, deduplication, aggregation, and embedded data quality checks - Use services such as S3, Lambda, EventBridge, Step Functions, SQS, SNS, Kinesis, Athena, Redshift, Lake Formation, CloudWatch, and IAM - Apply least-privilege IAM, encryption, secret management, and audit logging, and automate deployments via infrastructure as code and CI/CD - Monitor, troubleshoot, and optimize production pipelines for performance, reliability, scalability, and cost - Produce technical documentation including source-to-target mappings, runbooks, lineage, and operational procedures - Partner with product owners and stakeholders to refine epics, generate acceptance criteria, and track requirements through delivery and UAT
Requirements - 7+ years of relevant data engineering experience - Strong Python skills across modular design, testing, debugging, packaging, and performance optimization - Hands-on experience with AWS Glue, S3, the Glue Data Catalog, and broader AWS data, integration, security, and monitoring services - Practical experience ingesting, parsing, validating, transforming, and troubleshooting data, including schema drift and large files - Familiarity with data modeling, partitioning, metadata, lineage, and data quality practices - Experience with Git-based version control, automated testing, CI/CD pipelines, and infrastructure-as-code approaches - Strong analytical, documentation, and communication skills; experience presenting to both technical and non-technical audiences - Proficiency with JIRA and Confluence for requirements management - Bachelor's degree - Must be able to obtain and maintain a Public Trust clearance, including having lived in the United States for 3 of the past 5 years
Nice to have - SAFe Agile Certification - AWS certifications relevant to data engineering, architecture, or development - Healthcare IT experience and familiarity with HIPAA - Designing source-to-target mappings, canonical models, and integration patterns across heterogeneous providers - Partnering with data quality and governance teams on metrics, validation, profiling, and remediation - Managing SLAs and operational readiness processes - Experience with CMS programs, processes, and standards - Apache Spark or PySpark, Parquet, Avro, Iceberg, or other distributed processing and open table formats
Benefits and work setup - Flexible paid time off and hybrid working arrangements - Health coverage including medical, dental, vision, life, and disability insurance - Competitive compensation with a 401(k) including employer contributions - Training and development programs to build new skills and prepare for leadership roles