Remote job
Software Engineer, Data Infrastructure & Acquisition - Fresno, CA, USA
Job details
About this role
Role overview Take ownership of the data layer behind an AI team that builds audio datasets for next-generation text-to-speech products. The Fresno-based opening sits at the intersection of infrastructure, engineering, and research, focused on sourcing and ingesting training data at petabyte scale. It is a hands-on software engineering role inside a globally distributed, asynchronous organization.
Responsibilities - Discover new audio sources and wire them into the production ingestion pipeline - Operate and extend cloud infrastructure running on Google Cloud and managed with Terraform - Work with research scientists to shift the cost, throughput, and quality frontier for training datasets - Co-author the dataset roadmap with AI leadership to support consumer and enterprise offerings - Build automation that keeps the ingestion pipeline reliable as data volume grows
Requirements - BS, MS, or PhD in Computer Science or a closely related field - 5+ years of professional software development experience - Strong bash and Python scripting skills on Linux systems - Working knowledge of Docker and Infrastructure-as-Code, plus hands-on experience with a major cloud provider - Strong written and verbal communication skills - Ability to manage multiple tasks and adapt when priorities shift
Nice to have - Background with web crawlers and large-scale data processing workflows
Benefits and work setup - 100% distributed team with no physical office, built around asynchronous work - Colleagues with an entrepreneurial mindset who support calculated risk and ownership - Hands-off management style so engineers can focus on deep, high-impact work - US base salary range of $140,000 to $200,000, plus bonus and equity depending on experience - Friendly, laid-back atmosphere and a mission tied to accessibility, supporting users with dyslexia, ADD, low vision, concussions, autism, and other learning differences