Loading...
A zero-shot voice-cloning TTS model synthesises speech from a text script in the voice of a reference speaker; that generated audio is then fed into a separate lip-sync or talking-head model which drives realistic facial animation on a reference image or video, producing a complete avatar video without studio recording.
Source: https://arxiv.org/abs/2509.12831