发表机构
Computer Vision Center, Universitat Autònoma de Barcelona; Mohamed bin Zayed University of Artificial Intelligence; Jilin University; City University of Hong Kong (Dongguan)(巴塞罗那自治大学计算机视觉中心; 穆罕默德·本·扎耶德人工智能大学; 吉林大学; 香港城市大学(东莞))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有3D风格化无法独立控制颜色与纹理的问题,提出免训练框架DiDE,利用潜空间通道分区实现解耦注入,在Disen3D-Bench上优于基线。
AI 中文摘要
近期基于整流流(rectified flow)的图像到3D生成模型的进展,使得高保真3D资产生成成为可能。在此基础上,越来越多的研究工作利用这些强大的3D先验进行免训练风格化,将参考图像的视觉属性迁移到生成的3D资产上。然而,现有方法遵循“全有或全无”范式:颜色和纹理被联合迁移,缺乏独立控制它们的机制——我们将这一局限形式化为“解耦3D风格化(Disen3D)”。为解决此问题,我们提出DiDE,这是首个用于Disen3D的免训练框架。我们方法的关键观察在于:图像到3D模型的结构化潜空间相对于纹理是过完备的——纹理信息仅占据风格显著通道的一小部分子集,从而留下一个自由子空间用于独立编码颜色。DiDE利用这一特性,通过通道分区机制,将内容图像、纹理参考和颜色参考分别经由专用分支处理,并在每个自注意力层无干扰地组合两种风格信号,全程保持内容几何结构。在Disen3D-Bench(我们新收集的多参考基准)上的实验表明,DiDE在颜色保真度、纹理迁移和内容保持方面持续优于2D和3D风格化基线。
英文摘要
Recent advances in rectified flow-based image-to-3D generative models have enabled high-fidelity 3D asset generation. Building on this, a growing line of work has exploited these strong 3D priors for training-free stylization, transferring visual attributes from a reference image onto a generated 3D asset. However, existing methods enforce an all-or-nothing paradigm: color and texture are transferred jointly, with no mechanism to control them independently -- a limitation we formalize as Disentangled 3D Stylization(Disen3D). To address this, we propose DiDE, the first training-free framework for Disen3D. Key to our approach is the observation that the structured latent space of image-to-3D models is overcomplete with respect to texture: texture information occupies only a small subset of the style-significant channels, leaving a free subspace available for independent color encoding. DiDE exploits this via a channel partition mechanism that processes a content image, a texture reference, and a color reference through dedicated branches and composes both style signals interference-free at every self-attention layer, preserving content geometry throughout. Experiments on Disen3D-Bench, our newly collected multi-reference benchmark, show that DiDE consistently outperforms 2D and 3D stylization baselines in color fidelity, texture transfer, and content preservation.
CommentsAccepted to NeurIPS 2026