Remote job
Software Engineer, Data Infrastructure & Acquisition - Cupertino, CA, USA
Job details
About this role
Role overview A fully distributed text-to-speech company is seeking a Software Engineer in Cupertino to lead data infrastructure and acquisition for its AI training pipelines. The role blends infrastructure engineering with research collaboration to build high-quality, petabyte-scale audio datasets that fuel next-generation consumer and enterprise products. The organization operates 100% remotely with an asynchronous culture and no headquarters.
Responsibilities - Track down new audio data sources and route them through the ingestion pipeline. - Maintain and grow the cloud ingestion stack running on GCP and managed with Terraform. - Work alongside scientists to improve the cost, throughput, and quality balance of training data. - Shape the AI team's dataset strategy that backs upcoming product launches. - Juggle several concurrent initiatives and adjust priorities as the work evolves.
Requirements - BS, MS, or PhD in Computer Science or a related discipline. - 5+ years of professional software development experience. - Strong bash and Python scripting skills in Linux environments. - Comfort with Docker and Infrastructure-as-Code on at least one major cloud provider. - Effective written and verbal communication skills.
Nice to have - Background with web crawlers or large-scale data processing workflows.
Benefits and work setup - United States base salary range of $140,000 to $200,000, plus bonus and equity, depending on experience. - Fully remote, asynchronous team structure. - Entrepreneurial culture encouraging risk-taking and independent ownership. - Hands-off management style and friendly, laid-back working atmosphere. - Product with global reach that supports users with learning differences such as dyslexia, ADHD, low vision, concussions, and autism.