Remote job
Lead Data Engineer
Job details
About this role
Role overview
A Lead Data Engineer opportunity supporting large-scale federal data environments, focused on designing and operating cloud-native data platforms on AWS. The role pairs hands-on pipeline and platform engineering with technical leadership, including mentoring junior engineers, guiding architecture decisions, and leading modernization from legacy systems to cloud-native services.
Responsibilities
- Architect and maintain scalable data pipelines, ETL/ELT workflows, and data models using Python, PySpark, dbt, SQL (PostgreSQL), and AWS Glue. - Build AWS-native platforms leveraging Glue, EMR, MWAA (Airflow), Lambda, Step Functions, S3, Redshift, RDS, DMS, and CloudWatch, with CloudFormation for infrastructure as code. - Develop high-performance ingestion, transformation, and orchestration for structured and semi-structured data using Apache Iceberg, Parquet, ORC, and Avro. - Integrate relational and NoSQL sources (PostgreSQL, Oracle, Redshift, GraphDB, and others) into analytical platforms including Athena, Trino, Hive, and OpenSearch. - Support AI-enabled capabilities such as RAG pipelines and vector search using Amazon Bedrock and vector indexes, and migrate legacy systems (IBM DataStage, Hadoop, shell workflows) to cloud-native AWS services.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related technical field. - 15+ years of professional experience in data engineering, data architecture, or related disciplines. - U.S. citizenship and ability to obtain Public Trust clearance. - Strong hands-on experience with PySpark, Python, SQL/PostgreSQL, and dbt for large-scale ETL/ELT development and data modeling. - Deep familiarity with the AWS data and analytics services listed above and modern CI/CD practices using GitHub and Harness. - Experience working in Agile environments, with strong analytical, problem-solving, and communication skills.
Nice to have
- Background supporting financial regulators, capital markets, or other highly regulated environments. - Apache Iceberg and modern data lakehouse architecture experience. - Unstructured data processing, embeddings, vector search, and LLM-based data solutions. - Designing data architectures that handle both structured and unstructured data at scale.
Benefits and work setup
- Estimated salary range of approximately $150,000–$200,000 annually, dependent on experience and qualifications. - Remote work setup, with hybrid arrangements noted where applicable. - Medical, dental, and vision coverage plus life insurance and short/long-term disability. - 401(k) plan with 4% employer match and an employee assistance program. - Generous PTO, continuing education budget, and an annual wellness allowance. - Employee referral and performance-based bonus incentive programs.