Loading...
Given a reference image containing multiple people, separate audio streams per speaker, and a scene prompt, the pipeline generates a video where each character's lips and turn-taking match their own audio stream, driven by a large video diffusion backbone conditioned on per-person audio embeddings.
Source: https://arxiv.org/abs/2505.22647