Remote job
Software Engineer, Data Infrastructure & Acquisition - Shanghai, China
Job details
About this role
Role overview Shape the data foundation of an AI team that produces audio datasets for next-generation text-to-speech products. The Shanghai-based role combines infrastructure, engineering, and research collaboration to source, ingest, and refine training data at petabyte scale. It is a hands-on software engineering opening inside a globally distributed, asynchronous organization.
Responsibilities - Hunt for new audio data sources and route them into the production ingestion pipeline - Run and extend cloud infrastructure on Google Cloud, managed through Terraform - Partner with research scientists to move the cost, throughput, and quality frontier for training data - Work with AI leadership to define the dataset roadmap backing consumer and enterprise products - Harden and automate the ingestion pipeline as data volume and complexity grow
Requirements - Bachelor's, Master's, or PhD in Computer Science or a related discipline - 5+ years of professional software development experience - Proficiency in bash and Python scripting on Linux environments - Strong grasp of Docker and Infrastructure-as-Code, plus professional experience with a major cloud provider - Strong written and verbal communication skills - Ability to juggle multiple tasks and adapt to changing priorities
Nice to have - Experience with web crawlers and large-scale data processing workflows
Benefits and work setup - Fully distributed team with no central office and a deliberately asynchronous culture - Entrepreneurial peers who value initiative, intuition, and follow-through - Minimal managerial overhead so engineers can concentrate on building - Competitive compensation and the chance to ship a product used by millions of people - Mission tied to accessibility, supporting readers with dyslexia, ADD, low vision, concussions, autism, and other learning differences