Remote job
Senior Data Engineer
Job details
About this role
Role overview A senior-level Data/ML Engineering position on a Business Intelligence team, focused on designing the end-to-end data architecture that feeds a warehouse, LLM-driven applications, and AI-backed analytics. The work blends traditional large-scale data engineering with newer ML/LLM operations, supporting ingestion, transformation, semantic search, and conversational interfaces. The role is remote and targets candidates in Latin America, with in-person verification as part of the hiring process.
Responsibilities - Design, build, and maintain scalable pipelines that move data from diverse sources into centralized feature stores, training workflows, and real-time inference services. - Engineer retrieval workflows over unstructured data, including embeddings, indexing, and semantic search patterns that support RAG-style applications. - Develop lightweight analytics and dashboarding experiences that surface natural-language query capabilities and AI-generated insights. - Define processes for prompt engineering, agent orchestration, and model fine-tuning routines that power conversational interfaces. - Manage vector data stores and the indexing strategies that keep retrieval-augmented workflows fast and reliable. - Partner with data and business stakeholders to translate language-model use cases into scalable, production-ready solutions, and document all pipelines and model deployment routines.
Requirements - 8+ years of hands-on experience as a Data Engineer. - Strong proficiency in Python for transformation, manipulation, and large-scale processing. - Production experience with big data tooling such as Apache Spark, Hadoop, and Kafka for distributed and real-time workloads. - Demonstrated ability to design pipelines that ingest from RDBMS, JSON, API, and flat-file sources. - Advanced SQL and PL/SQL skills, deep BI and data-warehouse knowledge, plus cloud warehouse experience (Snowflake or Redshift). - Solid grasp of software engineering principles, comfort on Unix/Linux/Windows, Agile workflows, and version control in distributed environments.
Nice to have - Vector database experience (e.g., DataStax AstraDB) and LLM application frameworks such as LangChain or LlamaIndex, including prompt engineering, RAG, and orchestration. - Familiarity with open-source LLM ecosystems like Hugging Face Transformers, including fine-tuning and inference optimization. - MLOps tooling and CI/CD pipelines for model versioning and automated deployments.
Benefits and work setup - Fully remote work arrangement. - B2B employment with USD-denominated gross compensation. - Hardware provisioning, long-term stability, and a referral program. - Sponsorship of professional training, seminars, and conferences, plus company-supported English classes.