Job details
About this role
Role overview This position sits on the data side of an AI team supporting text-to-speech products that serve tens of millions of users across mobile, desktop, and browser platforms. The engineer will combine infrastructure, scripting, and research collaboration to expand a petabyte-scale dataset pipeline powering next-generation voice models. It is a fully remote role based in or around Bordeaux, France.
Responsibilities - Discover and integrate new audio data sources into the ingestion pipeline. - Maintain and grow GCP-based cloud infrastructure orchestrated with Terraform. - Work closely with research scientists to improve cost, throughput, and quality of training data. - Contribute to the dataset roadmap alongside AI team members and product leadership. - Ensure pipelines remain reliable and integrated with downstream model training operations.
Requirements - BS, MS, or PhD in Computer Science or a related discipline. - 5+ years of professional software engineering experience. - Strong bash and Python scripting skills in Linux environments. - Practical experience with Docker and Infrastructure-as-Code concepts. - Professional experience with at least one major cloud provider, ideally GCP. - Clear written and verbal communication, plus adaptability to shifting priorities.
Nice to have - Background building web crawlers or large-scale data processing systems.
Benefits and work setup - 100% remote, distributed-first culture with no physical office. - Asynchronous workflows with minimal management overhead. - Friendly, laid-back atmosphere in a fast-growing product environment. - Compensation range of $30,000–$100,000 USD per year, plus bonus and stock, depending on experience. - Opportunity to contribute to products that support people with dyslexia, ADHD, low vision, concussions, autism, and similar challenges.