Remote job
Software Engineer, Data Infrastructure & Acquisition - Yokohama, Japan
Job details
About this role
Role overview Join the data side of an AI team focused on building the audio datasets that power next-generation text-to-speech products. The Yokohama-based opening blends infrastructure engineering, scripting, and research collaboration to source, ingest, and refine training data at petabyte scale. It is a hands-on engineering role within a globally distributed, asynchronous organization.
Responsibilities - Scout and bring in new audio data sources, integrating them into the production ingestion pipeline - Operate and extend cloud infrastructure running on Google Cloud and managed through Terraform - Partner with research scientists to push the cost, throughput, and quality frontier for training data - Work with AI leadership to shape the dataset roadmap that underpins consumer and enterprise products - Build tooling and automation that keep the ingestion pipeline reliable as volume grows
Requirements - Bachelor's, Master's, or PhD in Computer Science or a related discipline - 5+ years of professional software development experience - Strong proficiency in bash and Python scripting on Linux systems - Working knowledge of Docker and Infrastructure-as-Code, plus hands-on experience with a major cloud provider - Strong written and verbal communication skills - Comfort juggling multiple priorities in a fast-changing environment
Nice to have - Background with web crawlers and large-scale data processing workflows
Benefits and work setup - 100% distributed team with no physical office, built around asynchronous work - Entrepreneurial colleagues who value initiative, risk-taking, and ownership - Hands-off management style that lets engineers focus on deep work - Competitive compensation on a product used by millions of people - A mission tied to accessibility, with downstream impact for readers who have dyslexia, ADD, low vision, concussions, autism, and other learning differences