Job details
About this role
Role overview Contribute to the data infrastructure that fuels a large consumer text-to-speech platform serving more than 50 million users. This Software Engineer role sits on the AI team's data side, owning ingestion pipelines that turn diverse audio sources into petabyte-scale training sets. You will split time between hands-on engineering and close collaboration with scientists and team leadership.
Responsibilities - Discover and integrate new sources of audio data into the team's ingestion pipeline. - Operate and evolve the GCP-based cloud infrastructure that underpins data ingestion, managed through Terraform. - Work with research scientists to push the cost, throughput, and quality frontier of training datasets. - Help define and execute the dataset roadmap that feeds next-generation consumer and enterprise products. - Ensure that pipeline improvements translate into richer data at lower cost per unit of quality.
Requirements - Degree (BS, MS, or PhD) in Computer Science or a related technical field. - 5+ years of industry experience writing production software. - Fluency in bash and Python scripting on Linux. - Hands-on experience with Docker and Infrastructure-as-Code, plus professional use of at least one major cloud provider (GCP experience preferred). - Ability to balance multiple priorities and adapt to shifting requirements. - Strong communication skills, both written and verbal.
Nice to have - Experience building web crawlers or working on large-scale data processing systems.
Benefits and work setup - Fully remote, distributed-first organization with no physical office. - Asynchronous culture that values focus time and minimal meetings. - Friendly, laid-back team environment with entrepreneurial energy. - Chance to ship a product that millions rely on, including users with learning differences such as dyslexia, ADHD, low vision, concussions, and autism. - Work at the intersection of AI and consumer audio technology.