Loading...
A text-to-image model generates a base image from a prompt, then a separate instruction-following inpaint model takes that image plus a mask and a text instruction and applies targeted changes to specific regions while leaving the rest intact.