Loading...
A multimodal LLM first reasons step-by-step about a video's objects, actions, and acoustic environment, producing a structured chain-of-thought plan that then steers a separate audio foundation model to generate foley, and the same MLLM can re-reason to guide targeted natural-language edits of specific sound elements.
Source: https://arxiv.org/abs/2506.21448