Remote job
Data Engineer
Job details
About this role
Role overview Prepare and operate client data pipelines that supply AI applications, with a focus on whether the underlying data is usable and trustworthy. You’ll assess source quality, connect to existing systems, build and test retrieval corpora, and deploy within client-controlled cloud environments.
Responsibilities - Audit client data sources, classify their quality, and identify what must be cleaned before AI systems can rely on them. - Connect pipelines to existing exports and legacy systems, including sources that clients cannot readily replace. - Build corpora for RAG applications and test their structure and coverage to expose gaps that could lead to unreliable answers. - Deploy and operate data systems in client cloud accounts, adapting to the environment provided. - Explain data quality limitations to clients and recommend practical next steps when issues block delivery. - Create an automated reporting agent that delivers recurring decision support without manual intervention.
Requirements - Ability to assess, clean, and prepare data for dependable use by AI applications. - Experience connecting pipelines to established data sources, including legacy enterprise systems. - Understanding of how corpus structure and quality affect retrieval and generated answers. - Ability to work within client-managed cloud environments, including AWS, Google Cloud, or Azure. - Clear communication skills for discussing data constraints and recommendations with client stakeholders. - Attention to data sovereignty and relevant privacy and regulatory requirements.