发表机构
CUHK; Tencent Hunyuan; Tsinghua University; Fudan University; Shanghai Innovation Institute; Nankai University(香港中文大学; 腾讯混元; 清华大学; 复旦大学; 上海创新研究院; 南开大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Tex-Zero证明无需3D资产,仅用2D图像构建训练数据即可训练高保真原生3D纹理生成模型,通过平面化图像与VAE/DiT架构实现精细纹理生成。
AI 中文摘要
原生3D纹理生成直接在3D空间中为给定几何体合成颜色,并以多视角参考图像为条件。人们普遍认为,训练此类模型需要大规模、高质量的真实3D资产数据,而获取这些数据一直是一个长期且具有挑战性的问题。在这项工作中,我们提出了Tex-Zero,证明了无需3D资产即可训练出高保真度的原生3D纹理生成框架。我们的关键观察是,只有高质量且精细的颜色信息对3D纹理训练至关重要,而所需的几何信息则不那么关键,可以手动构建而非从真实3D资产中获取。这一发现使得将大量高质量的2D图像转化为有效的3D纹理生成训练样本成为可能。具体而言,我们通过将每张图像表示为3D空间中的一个平面,并应用逐块随机旋转和聚合来构建复杂的几何结构,从而将高质量的2D图像转换为3D训练样本。利用这些构建的图像数据,我们训练了Tex-Zero VAE,尽管在训练过程中从未见过真实3D资产,它仍能高质量地重建它们。基于Tex-Zero VAE,我们同样仅使用构建的图像数据训练了Tex-Zero DiT,其中条件2D多视角图像被转换为3D空间中的平面,并由Tex-Zero VAE进行编码,从而减少了表示差距并提高了生成质量。大量实验表明,Tex-Zero仅使用图像作为训练数据即可生成具有精细细节的高保真3D纹理,为扩展3D纹理生成的数据范式提供了一个有前景的视角。
英文摘要
Native 3D texture generation synthesizes colors directly in 3D space for a given geometry, conditioned on multi-view reference images. It is generally believed that training such models requires large-scale, high-quality real 3D asset data, whose acquisition remains a long-standing and challenging problem. In this work, we propose Tex-Zero, demonstrating that a high-fidelity native 3D texture generation framework can be trained without 3D assets. Our key observation is that only high-quality and fine-grained color information is essential for 3D texture training, while the required geometric information is less critical and can be manually constructed rather than obtained from real 3D assets. This finding makes it possible to transform abundant, high-quality 2D images into effective training samples for 3D texture generation. Specifically, we convert high-quality 2D images into 3D training samples by representing each image as a plane in 3D space and applying patch-wise random rotations and aggregation to construct complex geometric structures. Using these constructed image data, we train the Tex-Zero VAE, which can reconstruct real 3D assets with high quality despite never observing them during training. Building upon the Tex-Zero VAE, we train the Tex-Zero DiT also exclusively on the constructed image data, where the conditioning 2D multi-view images are transformed into planes in 3D space and also encoded by the Tex-Zero VAE, thereby reducing the representation gap and improving generation quality. Extensive experiments show that Tex-Zero generates high-fidelity 3D textures with fine-grained details solely using images as training data, offering a promising perspective on the data paradigm for scaling 3D texture generation.
CommentsProject Page: https://github.com/wangjiangshan0725/Tex-Zero