Remote job
Research Scientist - Interactive Avatars
Job details
About this role
Role overview
Join an applied research team building the next generation of generative video agents that listen, speak, and react in real time. The work centers on interactive avatar technology: models which interpret user audio and video, then produce natural conversational responses with appropriate gaze, turn-taking, and body language. This is an end-to-end research role with a strong production orientation, ideal for someone who wants to push the frontier of multimodal diffusion while seeing their work ship into a live product.
Responsibilities
- Drive progress on dyadic interaction modeling, owning well-scoped research problems from hypothesis through evaluation - Advance perceptual capabilities of interactive agents, including parsing user audio/video and synthesizing appropriate reactive outputs - Post-train multimodal models that produce rich, lifelike two-way interactions from user inputs - Adapt diffusion architectures to novel conditioning signals such as conversational state, turn-taking cues, and listener feedback - Design and maintain evaluation frameworks and test suites that continuously track interaction quality - Collaborate with dataset teams to specify data needs and shape high-quality, task-relevant datasets - Run rigorous experiments and communicate findings that shape technical decisions across the team
Requirements
- Strong machine learning background with hands-on experience building diffusion models, ideally for video or avatar generation - Track record of publications at top venues such as CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, or SIGGRAPH in world models, video diffusion, or dyadic interaction, or equivalent demonstrated real-world impact - Experience moving research ideas from concept to working implementations, with motivation to reach production - Proficiency in PyTorch and the tooling required for large-scale model training - Clear written and verbal communication of hypotheses, experiments, and results
Nice to have
- Experience with real-time or streaming generation, including autoregressive video diffusion - Familiarity with distillation or other low-latency inference approaches - Audio-driven facial, gesture, or full-body motion modeling - Conversational modeling work such as turn-taking, backchanneling, or listener-response generation - Mentorship of junior researchers or students
Benefits and work setup
- Contribute to production-scale video foundation models within a fast-growing generative AI environment - Tackle hard research problems in scaling, stability, and controllability of human-centric video - Influence the direction of next-generation synthetic human technology - Join a highly technical, high-ownership culture focused on building and shipping - Engage with topics in AI safety, ethics, and security as core elements of the work - Work on a product trusted by a broad enterprise customer base across global industries