Stack a few reference photos of a person into a unified identity embedding, then generate that exact person in any scene, style, or action from text alone, with no per-subject training.
Properties
family
character-consistency
modality
images + text -> image
whyMultiModel
Identity fusion across multiple reference images needs a dedicated encoder and a merged embedding that a vanilla text-to-image model cannot represent.
steps
[object Object], [object Object], [object Object]
controls
Number and variety of reference photos; identity strength; text prompt and style tags.
exampleStack
PhotoMaker (SDXL) -> upscaler.
useCases
[Personalized avatars][Same person across marketing scenes][Fast persona prototyping]
pitfalls
Too few or too similar references collapse identity diversity; strong style tags can overwhelm the identity signal.