Remote job
Director, Text-to-Speech Synthesis Research
Job details
About this role
Role overview
This director-level research position leads the end-to-end Text-to-Speech program, spanning research strategy, technical direction, and the models that reach production. It is a hands-on leadership role that combines setting roadmap-level direction with staying directly engaged in architectures, experiments, training runs, and evaluation. The mandate is to advance the frontier of neural speech generation while converting those gains into deployable systems that move measurable production metrics.
Responsibilities
- Own the TTS research and model roadmap, choosing which technical bets are worth pursuing and recognizing when an approach should change or be retired - Drive advances across neural audio modeling, prosody and expressiveness, controllability, multilingual generation, voice identity and consistency, data and training strategy, post-training, and inference performance - Remain deeply technical: review research, challenge assumptions, design experiments, diagnose model failures, and tackle the highest-leverage problems hands-on - Build evaluation and benchmarking systems that explain *why* models improve, pairing automated metrics with human perceptual assessment - Lead a mix of individual contributors and tech-lead managers, hiring and developing both, holding a high technical bar, and setting direction across sub-teams - Partner with engineering and product leadership on ship-readiness and represent the program internally and externally
Requirements
- Deep expertise in modern TTS, speech generation, or audio generative modeling, with a hands-on track record of personally training and improving large-scale neural models - Command of the contemporary speech-generation stack and the open problems behind naturalness, expressiveness, controllability, robustness, voice consistency, and inference cost - Demonstrated ability to set research direction under genuine uncertainty, prioritizing experiments, allocating compute, and ending approaches that aren't working - Experience leading researchers and research engineers through technical leaders, developing tech-lead managers and directing sub-teams while remaining technically influential - A working style in which AI tools are the default mode of operation, with an earned view of what they still cannot do in speech research - Ability to make complex technical tradeoffs legible to product, engineering, and executive audiences
Nice to have
- TTS or generative-audio models deployed at meaningful production scale - Experience building or substantially scaling a high-performing AI research organization - Sophisticated evaluation systems for generative speech, including expressive or multilingual generation, voice cloning, or controllable generation - Recognized external contributions through publications, open-source work, patents, or invited talks in speech synthesis, neural audio codecs, speech language models, or multimodal models - Background in fast-moving startup or research environments that routinely take models from idea to production