Job details
About this role
Role overview This software engineering role focuses on building the data foundation for an AI team behind widely used text-to-speech products on mobile, desktop, and browser platforms. The work sits at the intersection of infrastructure, scripting, and applied research, producing petabyte-scale training datasets at low cost. It is a fully remote position based in or around Antwerp, Belgium.
Responsibilities - Source new audio datasets and route them into the ingestion pipeline. - Run and extend GCP-based infrastructure managed through Terraform. - Work with research scientists to optimize the cost/throughput/quality frontier. - Shape the dataset roadmap with the AI team and product leadership. - Keep ingestion flows tightly integrated with downstream model training.
Requirements - BS, MS, or PhD in Computer Science or a related field. - 5+ years of professional software development experience. - Strong bash and Python scripting skills on Linux. - Solid grasp of Docker and Infrastructure-as-Code patterns. - Hands-on experience with at least one major cloud provider, ideally GCP. - Strong communication skills and comfort adapting to evolving priorities.
Nice to have - Experience with web crawlers and large-scale data processing systems.
Benefits and work setup - Fully remote, distributed-first environment with no physical office. - Asynchronous culture with minimal management overhead. - Friendly, laid-back atmosphere oriented around impact and ownership. - Compensation range of $30,000–$100,000 USD per year, plus bonus and stock, depending on experience. - Opportunity to contribute to products that support people with dyslexia, ADHD, low vision, concussions, autism, and similar challenges.