Job details
About this role
Role overview Help build the large-scale data backbone behind a consumer text-to-speech product used by tens of millions of readers. As a Software Engineer on the data acquisition and infrastructure team, you will own pieces of an ingestion pipeline that delivers petabyte-scale training datasets at low cost. Expect close collaboration with research scientists to keep pushing what is possible in audio AI.
Responsibilities - Identify and bring in fresh sources of audio data for the ingestion pipeline. - Maintain and grow the cloud infrastructure underpinning ingestion, with GCP and Terraform as the current foundation. - Partner with scientists to improve the tradeoffs between cost, throughput, and dataset quality. - Help shape the dataset roadmap for future consumer and enterprise releases. - Translate pipeline improvements into better, cheaper data for downstream model training.
Requirements - BS, MS, or PhD in Computer Science or a closely related discipline. - 5+ years of professional software engineering experience. - Comfortable writing bash and Python scripts in Linux environments. - Solid experience with Docker and Infrastructure-as-Code practices on at least one major cloud provider (internally built on GCP). - Comfort prioritizing across competing tasks in a fast-moving environment. - Strong written and verbal communication skills.
Nice to have - Background with web crawlers and large-scale data processing pipelines.
Benefits and work setup - Fully remote, distributed-first culture with no central office. - Asynchronous working style designed to protect deep focus time. - Competitive compensation in the $140,000–$200,000 base range, plus bonus and equity based on experience. - Opportunity to ship product used by millions, including people with learning differences such as dyslexia, ADHD, low vision, concussions, and autism. - Work at the intersection of AI and audio in one of tech's fastest-growing sectors.