Loading...
A silent generated (or real) video plus an optional text caption is fed into a joint video-audio-text diffusion model that synthesizes semantically matched, temporally synchronized sound effects and ambience directly from visual motion.
Source: https://arxiv.org/abs/2412.15322