- family
- character-consistency
- modality
- reference images (subject + object) + text -> composed still -> animated video
- whyMultiModel
- UNO's DiT adapter only produces a single still frame from reference images; it has no temporal/motion prior. A separate video diffusion model with its own vision encoder is required to turn the consistent still into motion without re-drifting identity.
- steps
- [object Object], [object Object], [object Object]
- controls
- number/order of reference images, text prompt for scene, UnoPE position mapping (subject vs. object disambiguation), CLIP Vision reference averaging across multiple stills to stabilize identity, negative reference image to suppress unwanted traits
- exampleStack
- ByteDance UNO (FLUX.1 dev) -> CLIP Vision encode -> Wan 2.2 image-to-video
- useCases
[product/character mockups that need to move][outfit-consistent social clips from a single compose][game/animation asset previsualization]
- pitfalls
- UNO can still show attribute confusion between subject and object on complex multi-reference prompts; if the composed still has inconsistent lighting/crop, the video model's identity embedding destabilizes and drifts within a few seconds of generated motion