Drive a single portrait photo with speech by estimating realistic 3D motion coefficients (head pose and expression) from the audio, then rendering a lip-synced talking-head clip.
Properties
family
audio-driven
modality
portrait + audio -> talking-head video
whyMultiModel
Speech-to-motion mapping and photo-real rendering are separate problems; a dedicated motion-coefficient model feeds a renderer that a generator alone cannot.
steps
[object Object], [object Object]
controls
Source portrait; audio clip; expression and pose amplification; stillness of background.
exampleStack
SadTalker (or a commercial talking-head API) -> upscale.
useCases
[Narrator and explainer avatars][Localization and dubbing][Talking-photo gifts]
pitfalls
Extreme head turns break the single-image 3D assumption; fast speech outruns the expression model.