Job details
About this role
Role overview This is a software engineering role on a data infrastructure team that powers large-scale dataset creation for AI model training. The position sits at the intersection of cloud infrastructure, data engineering, and applied research, supporting the collection and ingestion of audio data for next-generation speech and language models. The work centers on building petabyte-scale pipelines that balance cost, throughput, and data quality.
Responsibilities - Identify and onboard new sources of audio data, integrating them into existing ingestion workflows - Operate and extend cloud infrastructure (GCP) managed with Terraform to support high-volume data pipelines - Partner with research scientists to improve the cost, throughput, and quality trade-offs of dataset production - Design and refine large-scale data processing systems that deliver richer training material at lower cost - Contribute to roadmap planning for dataset strategy in coordination with team leadership and broader product groups
Requirements - BS, MS, or PhD in Computer Science or a closely related field - At least 5 years of professional software development experience - Proficiency in bash and Python scripting within Linux environments - Hands-on experience with Docker and Infrastructure-as-Code practices - Professional experience with at least one major cloud provider, ideally GCP - Strong written and verbal communication skills, with the ability to manage multiple priorities in a fast-paced setting
Nice to have - Experience designing or operating web crawlers - Familiarity with large-scale data processing workflows
Benefits and work setup - Fully distributed, remote-first work environment - Entrepreneurial culture that supports initiative, experimentation, and ownership - Hands-off management style that emphasizes deep focus and asynchronous collaboration - Competitive compensation and the chance to shape products that reach millions of users - Opportunity to contribute to accessibility-focused tools that support readers with dyslexia, ADHD, low vision, and other learning differences