arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DirectUV:基于图像条件的UV纹理生成与表面感知位置编码

DirectUV: Image-Conditioned UV Texture Generation with Surface-Aware Positional Encoding

Jiantao Lin, Yingjie Xu, Mingzhi Sheng, Yangkai Wei, Hao Chen, Ying-Cong Chen

arXiv 2609.34651首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); The Hong Kong University of Science and Technology(香港科技大学(广州); 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DirectUV提出表面感知位置编码的UV纹理扩散框架,利用3D表面坐标替代UV网格位置编码,提升接缝连贯性与遮挡区域纹理质量。

AI 中文摘要

为3D网格生成高质量的UV纹理仍然具有挑战性。多视图投影管线存在遮挡和视图不一致的问题,而最近直接在UV空间中生成纹理的方法仍然依赖辅助模块来提供3D信息,使得注意力机制局限于UV网格位置而非底层表面几何。这种不匹配限制了接缝和不相连UV岛之间的连贯性。我们提出了DirectUV,一个基于图像条件的UV纹理扩散框架,在预训练图像VAE的潜在UV空间中运行,其中扩散Transformer根据单个输入图像和粗略UV图对UV潜在表示进行去噪。其核心是表面感知位置编码(SAPE),用通过UV到表面对应关系获得的每令牌3D表面坐标编码替代标准2D网格位置编码。由于位置编码定义了注意力使用的距离度量,SAPE使令牌能够根据从表面对应关系导出的3D位置邻近性而非UV网格距离进行交互,从而恢复接缝和不相连岛屿之间的连贯性。多级扩展进一步将不同的注意力头分配给同一潜在UV补丁的逐渐细化子划分,使模型能够在多个粒度上推理表面结构。实验表明,DirectUV比其他基线生成更清晰且全局更一致的纹理,在投影方法留下空隙或拉伸纹理的遮挡和视图未见区域中改进最大。

英文摘要

Generating high-quality UV textures for 3D meshes remains challenging. Multi-view projection pipelines suffer from occlusion and view inconsistency, and recent methods that generate textures directly in UV space still rely on auxiliary modules to supply 3D information, leaving the attention mechanism tied to UV-grid positions rather than to the underlying surface geometry. This mismatch limits coherence across seams and disconnected UV islands. We propose DirectUV, an image-conditioned UV texture diffusion framework that operates in the latent UV space of a pretrained image VAE, in which a Diffusion Transformer denoises the UV latent given a single input image and a coarse UV map. At its core, Surface-Aware Positional Encoding (SAPE) replaces the standard 2D-grid positional encoding with encodings derived from per-token 3D surface coordinates obtained via UV-to-surface correspondence. As positional encodings define the distance metric used by attention, SAPE enables tokens to interact according to 3D positional proximity derived from surface correspondence rather than UV-grid distance, restoring coherence across seams and disconnected islands. A multi-level extension further assigns different attention heads to progressively finer subdivisions of the same latent UV patch, allowing the model to reason about surface structure at multiple granularities. Experiments show that DirectUV produces sharper and more globally consistent textures than other baselines, with the largest improvements in occluded and view-unseen regions where projection-based methods leave gaps or stretched textures.

CommentsAccepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑