arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成前的观察:物体感知增强单视图3D重建

OPERA: Object Perception Enhances Single-view 3D Reconstruction

Y Huynh, Duc Thanh Nguyen, Mohamed Abdelrazek

arXiv 2607.18630首次发表:更新:

发表机构

Deakin University; Applied Artificial Intelligence Initiative; School of Information Technology(迪肯大学; 应用人工智能倡议组织; 信息技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究单视图3D物体重建问题,提出利用预训练感知模型提取的感知信号驱动重建的方法,该方法与模型无关且可即插即用,实验验证其能显著提升重建效果,为3D重建提供新思路。

AI 中文摘要

物体感知与重建之间的关系在人类视觉中已得到充分确立,但在计算机视觉中仍未得到充分探索。在本文中,我们证明了学习到的物体感知可以显著增强3D重建。针对具有挑战性的单视图3D物体重建任务,我们提出了一种方法,该方法利用从预训练的感知模型中提取的感知信号(捕获语义和几何信息)来从物体的单张图像驱动其重建。我们的方法与模型无关,可以以即插即用的方式集成到各种重建方法中。在基准数据集中使用两种最先进的单视图3D重建管道进行的实验表明,我们的方法实现了一致且显著的改进,验证了将感知纳入生成的有效性。我们对方法及其应用的各个方面进行了深入分析。我们的项目页面位于此https URL。

英文摘要

Single-view 3D reconstruction is a challenging task in computer vision due to information missing from the single input image. Generative model-based approaches can produce plausible 3D objects from a single image, thanks to data-driven priors learnt from rich and large-scale datasets. However, plausible generation does not guarantee fidelity to the geometry and appearance of the particular input object. Inspired by object perception in human vision, we propose OPERA, a framework that guides multi-view diffusion sampling with pretrained perception models through lightweight alignment modules. These modules are trained independently while both the generative and perception models remain frozen, allowing multiple signals to be combined at inference without joint fusion training. We evaluate OPERA on two single-view 3D reconstruction baselines using subsets of Google Scanned Objects and OmniObject3D. On the primary baseline, combined guidance reduces mean Chamfer Distance by 30.4\% and 24.8\%, respectively, relative to unguided reconstruction. We also compare our method with recent image-to-3D models. We provide in-depth analyses of the design choices and their effects across datasets and backbones. Our project page is at https://opera-3d.github.io/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑