Remote job
Software Engineer, Data Infrastructure & Acquisition
Job details
About this role
Role overview
A software engineering role on the data side of an applied AI team, owning the pipelines that collect and ingest audio data at petabyte scale. The work tightly couples infrastructure, engineering, and research to keep model training datasets rich, affordable, and high throughput. The team operates fully distributed with no central office.
Responsibilities
- Identify and onboard new sources of audio data, feeding them into the ingestion pipeline. - Operate and extend the cloud infrastructure that powers ingestion, currently built on GCP and managed with Terraform. - Partner with research scientists to push the cost, throughput, and quality frontier, delivering richer data at larger scale and lower cost for next-generation models. - Collaborate across the AI team and broader leadership to shape the dataset roadmap supporting consumer and enterprise products.
Requirements
- BS, MS, or PhD in Computer Science or a related field. - At least five years of industry experience in software development. - Proficiency with bash and Python scripting in Linux environments. - Strong grasp of Docker and Infrastructure-as-Code concepts, with professional experience on at least one major cloud provider, ideally GCP. - Ability to handle multiple tasks and adapt to shifting priorities. - Strong written and verbal communication skills.
Nice to have
- Experience with web crawlers and large-scale data processing workflows.
Benefits and work setup
Fast-growing environment with room to shape product direction, an entrepreneurial team that values initiative, hands-off management, competitive compensation, a friendly asynchronous culture, and the chance to build products that support millions of users including people with dyslexia, ADHD, low vision, and other learning differences. Fully distributed with no physical office.