arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10531cs.CV

利用测试时部分观测引导图像到三维生成

Guiding Image-to-3D Generation with Test-Time Partial Observations

Jerred Chen, Simon Weber, Ronald Clark

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出无需训练的测试时引导框架,利用射线一致观测似然将部分几何观测整合进预训练图像到三维模型,显著提升几何保真度与视觉质量。

中文摘要 AI 辅助

图像到三维模型可以从单张RGB图像生成视觉上引人注目的三维资产,但其几何形状往往仅受可用观测的松散约束,限制了它们在需要几何保真度的应用中的使用。然而,在许多现实场景中,物体的部分几何观测可能在测试时可用。我们引入了一个无需训练框架,用于将此类证据整合到预训练的图像到三维生成模型中,而无需重新训练或微调。为此,我们通过定义在模型占用表示上的射线一致观测似然来引导生成,结合表面占用和自由空间证据。应用于SAM 3D及其多视图扩展时,我们的方法在不同可观测性水平下显著提高了几何保真度以及视觉质量。我们的结果表明,预训练的图像到三维模型可以通过显式的测试时引导有效整合部分几何观测,补充其学习到的生成先验,而无需修改底层模型。

英文摘要

Image-to-3D models can generate visually compelling 3D assets from a single RGB image, but their geometry is often only loosely constrained by the available observations, limiting their use in applications that require geometric fidelity. In many real-world settings, however, partial geometric observations of the object may be available at test time. We introduce a training-free framework for incorporating such evidence into pretrained image-to-3D generative models without retraining or finetuning. To do this, we guide generation using a ray-consistent observation likelihood defined over the model's occupancy representation, combining surface occupancy and free-space evidence. Applied to SAM 3D and its multi-view extension, our approach substantially improves geometric fidelity across different levels of observability, as well as visual quality. Our results demonstrate that pretrained image-to-3D models can effectively integrate partial geometric observations through explicit test-time guidance, complementing their learned generative priors without modifying the underlying model.

发表机构

  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

↑