Job details
About this role
Role overview This hire will own parts of the data acquisition and infrastructure stack behind an AI-powered text-to-speech platform used globally on iOS, Android, Mac, and the web. The work combines cloud infrastructure engineering, scripting, and research partnership to build petabyte-scale datasets efficiently. It is a fully remote software engineering role based in or around Greensboro, NC.
Responsibilities - Hunt down new audio data sources and feed them into the ingestion pipeline. - Operate and extend cloud infrastructure that runs on GCP and is managed with Terraform. - Collaborate with scientists to improve the tradeoffs between cost, throughput, and data quality. - Contribute to dataset strategy alongside AI team members and senior leadership. - Support downstream model training with reliable, well-instrumented data flows.
Requirements - BS, MS, or PhD in Computer Science or a closely related field. - 5+ years of industry software development experience. - Proficiency with bash and Python on Linux systems. - Practical experience with Docker and Infrastructure-as-Code tooling. - Professional background with at least one major cloud provider (GCP preferred). - Strong written and verbal communication, with flexibility to handle shifting priorities.
Nice to have - Experience with web crawlers and large-scale data processing pipelines.
Benefits and work setup - Fully remote, distributed team with no central office. - Hands-off management style supporting deep focus and ownership. - Competitive base salary, bonus, and equity (US-based range of $140,000–$200,000 depending on experience). - Entrepreneurial, asynchronous culture oriented around impact. - Mission-driven product that supports readers with dyslexia, ADHD, low vision, concussions, autism, and related conditions.