← Back to jobs

Remote job

Senior Data Scientist, Audio

Data Scientist Full-time Permanent USA

Job details

$165K – $220K • Offers Equity • Offers Bonus Salary
USA Eligibility
Senior Experience
Full-time Employment

About this role

Role overview

A senior data science role focused on conversational audio, sitting at the intersection of research, engineering, and data operations. The work centers on understanding what makes real-world speech data difficult — accents, dialects, overlapping talk, code-switching, domain vocabulary, and varied acoustic conditions — and turning that understanding into measurable model gains. This is a hands-on, high-leverage position with latitude to define an area from first principles.

Responsibilities

- Characterize large-scale conversational audio corpora across languages, acoustic conditions, domains, speaker demographics, and quality, and identify gaps where representation is thin. - Design and operationalize active-learning loops that systematically decide which data to prioritize next based on expected model impact. - Architect human-in-the-loop workflows, tooling, and model-assisted steps that maximize the value of limited human attention. - Build curated benchmarks and methodologies that allow honest, representative claims about model quality across diverse real-world speech. - Convert scrappy proofs of concept into repeatable, documented, production-grade pipelines that other teams can run independently. - Partner closely with research and engineering to translate findings into model and product improvements.

Requirements

- Hands-on experience on real data pipelines and model-facing problems in data science, machine learning, or applied research. - Strong Python and general data tooling skills, with comfort building analysis, scoring, and automation directly. - Practical experience with data characterization, data selection, active learning, or similar prioritization problems. - Working familiarity with speech/audio or NLP models, including reasoning about output quality, confidence, and error modes. - A track record of turning ambiguous, messy data situations into measurable improvements. - A bias toward building reusable systems rather than one-off notebooks, plus strong communication skills for translating complex findings to non-specialist audiences. - Active, hands-on use of modern AI tools integrated into daily workflow.

Nice to have

- Direct experience with automatic speech recognition, text-to-speech, multilingual or code-switched audio data. - Experience with ensemble labeling, pseudo-labeling, or LLM-assisted annotation. - Familiarity with data provenance, PII and GDPR-aware pipelines, or model-improvement compliance. - Experience building custom or fine-tuned models for specific domains. - Comfort working alongside research and engineering on shared infrastructure.

Skills detected in the listing

PythonGDPRLLM
Detected Sep 22, 2026
Last verified Sep 22, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight