发表机构
Shandong University; University of Edinburgh(山东大学; 爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ScaleBlind利用2D基础图像生成先验,无需真实尺度信息即可从部分点云补全完整3D形状,通过多视角补全与跨模态融合实现新最先进性能。
AI 中文摘要
点云补全旨在从部分点云中推断出完整的3D形状,并作为重建、编辑和模拟等下游任务的基础构建模块。尽管近期取得了进展,现有的基于学习的方法在训练和测试时的归一化过程中往往隐式地依赖真实形状尺度(GT-scale)的访问,假设了在现实世界推理中不可用的特权信息。这一隐藏假设限制了实际部署,并且一旦移除oracle GT-scale线索,可能导致严重的补全伪影,例如过度或不足补全以及嵌套壳。我们观察到,最近的基础图像生成模型展现出对物体和几何的强理解能力,并能产生多视角一致的渲染,使其成为无GT-scale 3D补全的有前景的先验。受此启发,我们提出了ScaleBlind,一种新颖的框架,利用基于基础模型的图像补全直接从部分输入中恢复全局尺度,然后忠实地产生3D补全。具体来说,ScaleBlind从渲染的部分视角中“梦想”出完整的多视角外观,将推断出的缺失区域提升回3D以获得几何感知的粗略补全,并通过一个强大的跨模态融合网络与原始部分点云进一步细化。通过利用2D基础先验,我们的方法在推理时无需访问GT-scale信息。此外,它为2D生成先验与3D点云补全之间提供了原则性的桥梁。大量实验证明了我们框架的优越性,使ScaleBlind成为点云补全任务的新最先进方法。
英文摘要
Point cloud completion aims to infer a complete 3D shape from a partial point cloud and serves as a fundamental building block for downstream tasks such as reconstruction, editing, and simulation. Despite the recent progress, existing learning-based methods often implicitly rely on access to the ground-truth shape scale (GT-scale) during both training- and testing-time normalization, assuming privileged information that is unavailable in real-world inference. This hidden assumption limits practical deployment and can lead to severe completion artifacts, e.g., over- or under-completion and nested shells, once the oracle GT-scale cue is removed. We observe that the recent foundation image generation models exhibit a strong capability of understanding objects and geometries, and producing multi-view consistent renderings, making them promising priors for GT-scale-free 3D completion. Motivated by this insight, we propose ScaleBlind, a novel framework that leverages foundation-model-based image completion to recover global scale directly from partial inputs and then faithfully produces the 3D completion. Specifically, ScaleBlind dreams out complete multi-view appearances from rendered partial views, lifts the inferred missing regions back into 3D to obtain a geometry-aware coarse completion, and further refines it via a powerful cross-modal fusion network with the original partial point cloud. By harnessing 2D foundation priors, our method eliminates the need for accessing GT-scale information at inference. Moreover, it provides a principled bridge between 2D generative priors and 3D point cloud completion. Extensive experiments demonstrate the superiority of our framework, making ScaleBlind the new state-of-the-art for the point cloud completion task.