Loading...
ByteDance's Bernini splits instruction-based video generation and editing into two models: an MLLM planner that predicts the target semantic representation in ViT embedding space, and a DiT diffusion renderer that synthesizes pixels conditioned on that plan plus source-video features, keeping unedited regions stable and preserving subject identity.
Source: https://arxiv.org/abs/2605.22344