Remote job
Research Staff, Voice AI Foundations [Remote]
Job details
About this role
Role overview Research and develop latent-space approaches to make voice AI more robust, expressive, and practical to train and deploy at scale. The work may span audio compression, generative speech, representation learning, synthetic data, and efficient model training and inference. The research aims to support systems that understand varied speakers and environments and produce natural responses for conversation or task completion.
Responsibilities - Develop neural audio codecs that compress audio efficiently while preserving reconstruction quality across diverse material. - Explore controllable speech generation, including varied speaking styles, emotional expression, multiple speakers, and noisy conditions. - Build representations that distinguish factors such as speaker, spoken content, style, environment, and channel. - Investigate recombining latent representations to expand training data and support larger-scale audio model development. - Design model architectures, training methods, and inference algorithms with computational and hardware costs in mind.
Requirements - Comfort experimenting with AI tools and incorporating them into research workflows. - Ability to adapt as research priorities and tools evolve, and to learn continuously in a fast-changing environment.
Benefits and work setup The listing describes the role as remote.