- family
- character-consistency
- modality
- single face photo + text (+ optional pose video / style LoRA) -> identity-locked video
- whyMultiModel
- Identity extraction and video synthesis are handled by two separate networks: a dedicated face-recognition encoder produces the identity embedding, and a large pretrained video DiT (untouched in its core weights) consumes it; optional third-party pose control (VACE) or LoRA further compose in as independent modules.
- steps
- [object Object], [object Object], [object Object]
- controls
- single reference face image, text prompt, optional driving pose video (VACE), optional LoRA for stylization, denoising-strength adjustment for the experimental face-swap mode
- exampleStack
- antelopev2 face encoder -> Stand-In adapter + Wan 2.1-14B text-to-video -> VACE pose control (optional)
- useCases
[identity-locked talking/acting clips from one photo][face-swap-in-video variants][stylized (cartoon/anime) identity-consistent video via community LoRA composition]
- pitfalls
- trained only on real-person data, so cartoon/object generalization is a secondary, less-tested mode; combining pose control (VACE) with identity locking simultaneously can still trade off motion naturalness against strict facial consistency