Job details
About this role
Role overview Help power the data backbone of a text-to-speech product serving tens of millions of readers. This role owns audio data acquisition and the cloud infrastructure behind it, enabling petabyte-scale training datasets for consumer and enterprise voice models. You will work remotely alongside scientists and engineers, blending infrastructure-as-code work with research-driven exploration.
Responsibilities - Find and onboard fresh audio sources into the ingestion pipeline using a scrappy, experimental mindset. - Operate and expand ingestion infrastructure running on GCP, managed through Terraform. - Collaborate with research scientists to optimize the trade-offs between cost, throughput, and dataset quality. - Help shape the dataset roadmap in coordination with AI team leadership to fuel upcoming product launches. - Build and refine tooling that supports large-scale data processing workflows.
Requirements - BS, MS, or PhD in Computer Science or a closely related field. - 5+ years of industry software engineering experience. - Strong scripting skills in bash and Python on Linux systems. - Solid grasp of Docker and Infrastructure-as-Code patterns. - Production experience with a major cloud provider, ideally GCP. - Strong written and verbal communication plus flexibility to handle shifting priorities.
Nice to have - Experience building web crawlers or managing large-scale data pipelines.
Benefits and work setup - 100% remote, asynchronous culture with no physical office. - United States base salary range of $140,000-$200,000, plus bonus and equity, depending on experience. - Mission-driven product that supports people with learning differences such as dyslexia, ADHD, low vision, concussions, and autism. - Entrepreneurial team environment with a hands-off management style.