Job details
About this role
Role overview Join the team building the data pipeline behind a text-to-speech platform used by more than 50 million people. The role sits at the intersection of cloud infrastructure engineering and large-scale data acquisition, enabling petabyte-scale training datasets for new voice models. You will work in a fully remote, asynchronous environment with a team drawn from major tech companies and research programs.
Responsibilities - Hunt down new audio data sources and feed them into the ingestion pipeline with a pragmatic, exploratory approach. - Operate and grow cloud infrastructure on GCP, managed through Terraform, that powers continuous dataset ingestion. - Partner with research scientists to push the cost-versus-quality frontier, producing richer data at greater scale and lower cost. - Help define the dataset roadmap alongside AI team leadership to support next-generation products. - Build and maintain tooling that keeps large-scale data processing workflows reliable and efficient.
Requirements - BS, MS, or PhD in Computer Science or a closely related field. - 5+ years of industry experience in software development. - Strong scripting skills in bash and Python on Linux systems. - Solid grasp of Docker and Infrastructure-as-Code patterns. - Professional experience with at least one major cloud provider, ideally GCP. - Strong written and verbal communication, with the ability to juggle shifting priorities.
Nice to have - Experience building web crawlers or operating large-scale data pipelines.
Benefits and work setup - Fully remote, asynchronous culture with no office. - United States base salary range of $140,000-$200,000, plus bonus and equity, depending on experience. - Mission-driven product supporting users with dyslexia, ADHD, low vision, concussions, autism, and other learning differences. - Entrepreneurial team with a hands-off management style that emphasizes autonomy and ownership.