Loading...
A text-to-image model generates one or more keyframe stills that lock composition, character, and lighting; those stills are then fed into an image-to-video model to produce motion. The i2v model inherits the keyframe's identity rather than hallucinating from text alone, giving tighter control over subject appearance and scene layout.