Loading...
A first model generates an untextured 3D geometry (mesh / TSDF) from a single image or text prompt; a second, separate diffusion model synthesises high-resolution texture maps conditioned on the produced geometry, yielding a fully textured, PBR-ready 3D asset.
Source: https://arxiv.org/abs/2505.07747