Loading...
LTX-2's dedicated audio-to-video pipeline generates a synchronized base clip directly conditioned on an input audio file (not text), jointly denoising audio and video in one diffusion pass, then a separate spatial-upscaler checkpoint sharpens the result in a second chained pass.