Remote job
Software Engineer, Data Infrastructure & Acquisition - Madison, WI, USA
Job details
About this role
Role overview This role sits on the data side of an AI team building large-scale datasets for a popular text-to-speech platform used globally across iOS, Android, Mac, Chrome, and the web. The engineer is responsible for the end-to-end pipeline that turns raw audio sources into training material at petabyte scale, integrating infrastructure, engineering, and research work in a tight loop.
Responsibilities - Identify and bring in fresh sources of audio data, feeding them into the ingestion pipeline. - Run and evolve the cloud infrastructure that hosts the pipeline, currently built on GCP with Terraform. - Collaborate with scientists to improve the trade-offs between cost, throughput, and data quality for upcoming model generations. - Help shape the AI team's dataset roadmap in partnership with leadership, with an eye toward both consumer and enterprise products. - Juggle shifting priorities across research, infrastructure, and product timelines.
Requirements - BS, MS, or PhD in Computer Science or a related technical discipline. - 5+ years of industry software engineering experience. - Strong command of bash and Python on Linux systems. - Practical experience with Docker, infrastructure-as-code patterns, and at least one major cloud provider (GCP preferred). - Clear written and verbal communication, including the ability to explain trade-offs to both engineers and researchers.
Nice to have - Background designing web crawlers or other high-throughput data acquisition systems.
Benefits and work setup - 100% remote, distributed organization; asynchronous collaboration is the default. - Entrepreneurial culture with minimal management overhead and a laid-back atmosphere. - US base salary range of $140,000–$200,000, plus bonus and equity, depending on experience. - Opportunity to ship products that materially help people with learning differences and other reading barriers.