Remote job
Software Engineer, Data Infrastructure & Acquisition - Taipei, Taiwan
Job details
About this role
Role overview Help build the data backbone of an AI team creating audio datasets for next-generation text-to-speech products. The Taipei-based opening sits at the intersection of infrastructure, engineering, and research, focused on sourcing and ingesting training data at petabyte scale. It is a hands-on software engineering role inside a globally distributed, asynchronous organization.
Responsibilities - Track down new audio sources and feed them into the production ingestion pipeline - Run and extend the team's cloud setup on Google Cloud, managed through Terraform - Work alongside research scientists to improve the cost, throughput, and quality of training datasets - Co-author the dataset roadmap with AI leadership in support of consumer and enterprise products - Automate and harden the ingestion pipeline so it keeps pace with growing data volumes
Requirements - BS, MS, or PhD in Computer Science or a closely related field - 5+ years of professional software engineering experience - Fluent in bash and Python scripting on Linux environments - Solid grasp of Docker and Infrastructure-as-Code, with hands-on experience on a major cloud platform - Clear, confident written and verbal communication skills - Ability to handle shifting priorities across multiple workstreams
Nice to have - Experience building or operating web crawlers and large-scale data processing workflows
Benefits and work setup - Fully remote team with no central office and a strong asynchronous culture - Colleagues with an entrepreneurial mindset who back risk-taking, intuition, and hustle - Minimal managerial overhead so engineers can stay focused on building - Competitive pay and equity on a product serving a very large user base - Work that supports people with dyslexia, ADD, low vision, concussions, autism, and other learning differences through accessible reading tools