发表机构
City University of Hong Kong (Dongguan); SB Intuitions Corp.; Universitat Autònoma de Barcelona(香港城市大学(东莞); SB Intuitions 公司; 巴塞罗那自治大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对图像到三维生成中的细节衰减问题,提出无需训练的推理时框架BTC3D,利用混合瓦片嵌入和动态条件调度增强细节保留,提升纹理质量与视觉保真度。
AI 中文摘要
近年来,基于扩散的流程在图像到三维合成方面取得了显著进展。然而,生成高保真细节仍然具有挑战性,尤其是当输入图像包含丰富细节时。现有方法通常依赖全局编码的条件特征,这会压缩空间信息并限制模型再现细粒度细节的能力。这种常见设计往往导致我们称之为细节衰减的现象。此外,提高图像到三维合成质量通常需要对大型扩散模型进行重新训练或微调,这计算成本高昂,且对于复杂的三维流程而言不切实际。在这项工作中,我们提出了用于图像到三维生成的混合瓦片条件化(BTC3D),这是一种无需训练、仅在推理时运行的框架,可增强图像到三维扩散流程中的细粒度细节保留。为了缓解细节衰减,我们首先研究了图像到三维模型中的图像特征可加性。基于这一特性,我们引入了一种混合瓦片嵌入,从分割的图像区域块中提取局部条件信号,使扩散模型能够更好地保留细粒度的视觉细节。为了稳定地整合全局和局部条件引导,我们提出了一种动态条件化调度,在扩散的后期低噪声阶段逐步增加瓦片级条件化的影响。我们提出的方法BTC3D完全在推理时运行,可以无缝集成到现有的图像到三维扩散流程中。实验结果表明,所提出的方法显著提高了基础模型的纹理质量和视觉保真度,同时以无需训练的方式保持了全局结构一致性。
英文摘要
Recent diffusion-based pipelines have achieved promising progress in image-to-3D synthesis. However, generating high-fidelity details remains challenging, especially when the input image contains rich details. Existing approaches often rely on globally encoded conditioning features, which compress spatial information and limit the model to reproduce fine-grained details. This common design often leads to a phenomenon we term detail attenuation. Moreover, improving image-to-3D synthesis quality typically requires retraining or fine-tuning large diffusion models, which can be computationally expensive and impractical for complex 3D pipelines. In this work, we present Blended Tile Conditioning for image-to-3D generation (BTC3D), a training-free inference time framework that enhances fine-grained detail preservation in image-to-3D diffusion pipelines. To alleviate detail attenuation, we first examine the image feature additivity in image-to-3D models. Based on this property, we introduce a blended tile embedding that extracts local conditioning signals from split image regional patches, allowing the diffusion model to better preserve fine-grained visual details. To integrate the global and local conditioning guidance stably, we propose a dynamic conditioning schedule that gradually increases the influence of tile-level conditioning during later low-noise stages of diffusion. Our proposed method BTC3D operates entirely at inference time and can be seamlessly integrated into existing image-to-3D diffusion pipelines. Experimental results demonstrate that the proposed approach significantly improves texture quality and visual fidelity of the base model while maintaining global structural consistency in a training-free manner.