Remote job
Machine Learning / Data Engineer
Job details
About this role
Role overview
This is a senior, hands-on individual contributor role focused on safeguarding enterprise data as it flows through connectors, processing stages, and a sanitization layer. The core mission is twofold: ensuring data quality at every step and building the machine-learning pipelines that detect and reliably replace sensitive entities — including personal, health, company-identifiable, and financial information — in both text and image documents. The position is fully remote for candidates based in Brazil or Colombia and pairs ML engineering with a quality-assurance mindset.
Responsibilities
- Assess enterprise data quality across connectors, evaluating topic coherence, domain depth, completeness, and consistency from raw through processed to sanitized stages. - Design and automate validation suites — schema checks, drift detection, and reconciliation — and surface quality issues in ways engineering and product teams can act on. - Train, evaluate, and maintain ML models (NER and other approaches) that detect sensitive entities across text and image-based documents such as scans, invoices, and presentations. - Build replacement pipelines that consistently substitute detected entities, so a given entity always maps to the same replacement across every file in a corpus. - Construct adversarial test sets covering edge cases, obfuscated identifiers, multilingual entities, OCR noise, and formats designed to evade detectors; measure precision, recall, leak rates, and replacement consistency per data class. - Implement CI/CD regression gates so pipeline changes cannot ship without passing data-quality and sensitive-data checks, and run sampling-based human-in-the-loop audits to maintain a compliance evidence trail.
Requirements
- Around 4–5 years of hands-on machine learning experience, with ML as a primary background. - Strong Python for ML development and data validation, using tools such as pytest, Great Expectations, or Pandera, plus solid SQL. - Familiarity with sensitive-data categories and relevant standards — HIPAA Safe Harbor for PHI, GDPR/LGPD for PII, PCI DSS for cardholder data, and confidentiality or NDA obligations for company information. - Experience building and evaluating NER or ML-based detection systems, including constructing labeled evaluation sets, computing precision and recall, and handling non-determinism. - Understanding of re-identification risk and how combined details can reveal an organization or individual. - A skeptical, detail-oriented approach that assumes systems are flawed until proven otherwise, with comfort operating in ambiguity and a fast-moving environment.
Nice to have
- Computer vision and OCR experience, especially building and evaluating document pipelines for contracts, statements, invoices, and presentations. - Experience handling financial or healthcare data and the privacy requirements specific to those industries. - Prior startup experience. - Auditing LLM or VLM outputs. - Synthetic generation of PII, PHI, company, and financial records. - Familiarity with the GCP data stack (BigQuery, GCS, Cloud Run jobs) and CI/CD integration. - Experience with multi-tenant enterprise data under strict confidentiality requirements, including compliance reporting or working with auditors.
Benefits and work setup
- Fully remote position based in Brazil or Colombia. - Work on problems at the frontier of agentic AI applied to real-world enterprise workflows. - Opportunity to contribute to expert datasets, RL environments, and first-of-a-kind benchmarks, with the chance to present work at top-tier ML conferences. - Collaboration with engineers who have deep AI experience from leading technology companies. - Startup-style pace with high ownership and impact.