arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

单张图像到纹理化3D物体生成的频域方法:从理论到流水线

Single Image to Textured 3D Object Generation in Frequency Domain: From Theory to Pipeline

Qisen Wang, Yifan Zhao, Jia Li

arXiv 2609.07085首次发表:更新:

发表机构

Beihang University; State Key Laboratory of Virtual Reality Technology and Systems(北京航空航天大学; 虚拟现实技术与系统国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对单视图3D重建中颜色偏差、视图不一致和高频细节缺失问题,提出频域混合优化框架及Morpheus3D流水线,利用高通2D先验增强3D先验,显著提升生成质量。

AI 中文摘要

单视图3D重建,也称为图像到3D,由于信息的极度缺乏,是一项持续具有挑战性的任务。近年来,在大规模数据集上预训练的扩散模型作为2D先验被用于解决这一不适定任务,但存在颜色偏差和视图不一致的问题,这些问题可以通过使用带有3D标注数据微调的扩散模型作为3D先验来抑制。然而,3D先验缺乏高频细节,这无法通过在空间域中直接与2D先验互补来解决,因为这会引入错误的低频2D先验指导。在本文中,我们从频率角度重新审视不同扩散先验的特性。基于我们的观察,我们从理论上提出了一个统一的框架,用于在频域中使用多个扩散先验进行混合优化。在此框架下,我们进一步提出了Morpheus3D,一个从任意野外单张无姿态图像生成3D物体的流水线。Morpheus3D通过高通图像提示2D先验指导增强3D先验,以重建高质量的3D物体,同时有效抑制视图不一致、低频颜色偏差和高频缺失问题。在公开数据集和我们收集的具有复杂纹理的数据集上的定量和定性实验表明,我们的方法在生成质量上表现出显著改进。

英文摘要

Single-view 3D reconstruction, also known as image-to-3D, is a persistently challenging task due to the extreme lack of information. Recently, diffusion models pre-trained on large-scale datasets served as 2D priors are used to solve the ill-posed task but suffer from color deviation and view inconsistency, which can be curbed by using diffusion models fine-tuned with 3D annotated data served as 3D priors. However, 3D priors lack high-frequency details, which cannot be solved by direct complementation with 2D priors in spatial domain for introducing erroneous low-frequency 2D prior guidance. In this paper, we revisit the characteristics of different diffusion priors from the frequency perspective. Based on our observations, we theoretically present a unified framework of hybrid optimization using multiple diffusion priors in frequency domain. Under this framework, we further propose Morpheus3D, a pipeline of 3D object generation from any single unposed image in the wild. Morpheus3D enhances 3D prior with high-pass image-prompt 2D prior guidance to reconstruct high-quality 3D objects while effectively suppressing view inconsistency, low-frequency color deviation, and high-frequency lacking problems. Both quantitative and qualitative experiments on the public and our collected datasets with complex textures show that our method exhibits significant improvements in generation quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑