Remote job
Staff Research Scientist- Interactive Avatars
Job details
About this role
Role overview Join a large research and engineering organization focused on generative AI for video, working at the frontier of avatar-centric, interactive video diffusion models. The role centers on building AI video agents capable of listening, responding, and reacting naturally in real-time conversations, including turn-taking, gaze, facial expressions, and body language. It combines hands-on research with leadership, owning roadmap direction and translating breakthroughs into product capabilities.
Responsibilities - Set the research roadmap for dyadic interaction modeling, balancing long-term bets with near-term product impact. - Advance the state of the art in the perceptual layer of interactive agents, including understanding user audio and video and generating contextually appropriate reactions. - Post-train multimodal models to produce rich, natural dyadic interactions from user audio and video inputs. - Adapt diffusion models to new conditioning signals such as conversational state, turn-taking, and listener cues. - Build evaluation frameworks and test suites for continuous tracking of interaction quality. - Partner with data teams to define data needs and shape high-quality datasets, while mentoring researchers and driving cross-functional technical decisions.
Requirements - Deep machine learning expertise with extensive hands-on experience in diffusion models, ideally for video or avatar generation. - Strong publication record at top-tier venues (e.g., CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, SIGGRAPH) in world models, dyadic interaction, video diffusion, or equivalent demonstrated impact. - Experience leading a small team of researchers and mentoring junior members. - Track record of moving research from idea to production. - Proficiency in PyTorch and modern ML tooling for large-scale training. - Clear communication of hypotheses, experiments, and results, with the ability to shape direction across teams.
Nice to have - Experience with real-time or streaming generation, including autoregressive video diffusion. - Distillation or other techniques for low-latency inference. - Audio-driven facial, gesture, or full-body motion modeling. - Conversational modeling such as turn-taking, backchanneling, or listener-response generation.
Benefits and work setup - Competitive compensation. - Hybrid work setting with offices in several major European cities, or remote within Europe. - 25 days of annual leave plus public holidays. - Company culture described as built around doing focused, high-velocity work, with regular planning and social events at regional hubs, plus additional benefits depending on location.