Job details
About this role
Role overview Join the data team powering a text-to-speech platform used by over 50 million people worldwide. The work centers on collecting and processing audio datasets at petabyte scale for training next-generation voice models, blending cloud infrastructure engineering with research-driven data acquisition. The team operates fully remotely with an asynchronous culture and an emphasis on entrepreneurial execution.
Responsibilities - Scout and onboard new audio data sources into the ingestion pipeline, taking a pragmatic, hands-on approach to discovery. - Operate and extend cloud infrastructure that powers ingestion, currently built on GCP and managed through Terraform. - Partner closely with research scientists to push the cost, throughput, and quality frontier, delivering richer datasets at greater scale and lower cost. - Contribute to the dataset roadmap alongside AI leadership to support upcoming consumer and enterprise products. - Maintain and improve tooling for large-scale data workflows in Linux environments.
Requirements - BS, MS, or PhD in Computer Science or a related discipline. - 5+ years of professional software development experience. - Strong scripting ability in bash and Python within Linux environments. - Hands-on proficiency with Docker and Infrastructure-as-Code patterns. - Professional experience with at least one major cloud provider (GCP preferred). - Clear written and verbal communication skills and comfort juggling shifting priorities.
Nice to have - Background building web crawlers or operating large-scale data processing pipelines.
Benefits and work setup - Fully remote, asynchronous working environment with no central office. - Compensation range for Finland-based hires: 30,000-100,000 USD per year plus bonus and stock, depending on experience. - Entrepreneurial culture that encourages risk-taking and initiative. - Mission-driven work supporting users with dyslexia, ADHD, low vision, concussions, autism, and other learning differences.