arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UltraTex:释放2K多视图扩散用于3D纹理生成

UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing

Yibo Zhang, Ze Yuan, Nan Cao, Li Zhang, Yan-Pei Cao, Yuan-Chen Guo, Rui Ma

arXiv 2609.23169首次发表:更新:

发表机构

Jilin University; Shanghai Innovation Institute; The University of Hong Kong; Tongji University; Fudan University; VAST(吉林大学; 上海创新研究院; 香港大学; 同济大学; 复旦大学; VAST)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

UltraTex提出高效端到端框架,通过背景令牌丢弃和块稀疏注意力解决冗余,实现2K分辨率多视图扩散3D纹理生成,大幅提升效率并保持细节质量。

AI 中文摘要

高质量纹理生成对于创建逼真且可投入生产的3D资产至关重要。最近的多视图扩散方法在图像引导的3D纹理生成方面显示出有前景的结果,但它们通常受限于512或768的低操作分辨率,这使得难以从高分辨率参考图像中保留高频细节。将该范式扩展到2048分辨率在计算上是不可行的,因为统一的多视图序列超过212K个令牌,并导致过度的内存和延迟。在本文中,我们提出了UltraTex,一个用于基于高分辨率多视图扩散的3D纹理生成的高效端到端框架。我们的关键观察是,以对象为中心的多视图渲染包含两个主要的冗余来源:背景引起的序列冗余和前景内的稀疏令牌交互。为了解决这些问题,我们引入了背景令牌丢弃(Background Token Dropping),在DiT主干之前移除背景令牌,以及块稀疏注意力(Block-Sparse Attention),减少对保留的前景序列的注意力计算。为了实现仅前景的高效推理同时避免重建伪影,我们进一步设计了前景感知VAE解码(Foreground-Aware VAE Decoding),以确保最终高分辨率视图的质量。为了满足2K分辨率多视图扩散训练的高数据需求,我们构建了G-buffer TexVerse,一个大规模、超高分辨率的多视图渲染数据集,覆盖超过268,000个3D资产。大量实验表明,UltraTex生成具有丰富细粒度细节的视觉逼真纹理,同时大幅提高效率,在数据集中的常见样本上,相对于基线实现了$20.6\ imes$--$91.1\ imes$的训练加速和$22.3\ imes$--$74.6\ imes$的端到端推理加速。代码和数据位于此https URL。

英文摘要

High-quality texture generation is essential for creating realistic and production-ready 3D assets. Recent multi-view diffusion methods have shown promising results for image-guided 3D texturing, but they are typically constrained to low operating resolutions such as 512 or 768, making it difficult to preserve high-frequency details from high-resolution reference images. Scaling this paradigm to 2048 resolution is computationally prohibitive, as the unified multi-view sequence exceeds 212K tokens and incurs excessive memory and latency. In this paper, we present UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing. Our key observation is that object-centric multi-view renderings contain two major sources of redundancy: background-induced sequence redundancy and sparse token interactions within the foreground. To address them, we introduce Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequence. To enable efficient foreground-only inference while avoiding reconstruction artifacts, we further design Foreground-Aware VAE Decoding to ensure the quality of the final high-resolution views. To satisfy the demanding data requirements of 2K-resolution multi-view diffusion training, we construct G-buffer TexVerse, a large-scale, ultra-high-resolution multi-view rendering dataset covering over 268,000 3D assets. Extensive experiments show that UltraTex generates visually faithful textures with rich fine-grained details, while substantially improving efficiency, achieving $20.6\times$--$91.1\times$ training speedup and $22.3\times$--$74.6\times$ end-to-end inference speedup over the baseline on common samples in our dataset. Code and data is at https://yiboz2001.github.io/UltraTex.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑