发表机构
Massachusetts Institute of Technology; National University of Singapore(麻省理工学院; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Matisse是一个无需训练的框架,利用预训练生成式三维模型的证据来估计不确定性并指导主动重建和关键帧选择,在多个数据集上显著降低了Chamfer距离并提升了速度。
AI 中文摘要
一个三维重建系统如何在有限的计算预算下,从部分视角中获取并保留有用信息以理解场景的几何结构?现有的主动视角获取方法通常估计观测到的或实例化的几何上的不确定性,这限制了它们推理未见结构的能力,而长时程重建方法往往保留冗余的观测。我们提出了Matisse,一个无需训练的框架,通过利用预训练生成式三维模型提供的证据,统一了主动重建和关键帧选择。Matisse从与三维潜在标记相关联的交叉注意力证据中估计证据不确定性,并推导出证据信息增益,以基于后验熵的预期减少来指导视角获取和关键帧选择。Matisse通过遮挡感知、对象平衡的聚合支持多对象场景,并通过中间潜在变量传播不确定性,以避免规划过程中的完整重建。相对于每个数据集上的最佳基线,Matisse在GSO30、YCB-V和Replica上将Chamfer距离分别减少了12.7%、3.8%和9.2%,并在GSO30上使用相同的重建后端实现了比最佳主动重建基线快1.50倍的端到端加速。在GSO30长时程重建的关键帧选择实验中,Matisse使用14%的输入视图即可达到与Stream3D相当的Chamfer距离。
英文摘要
How can a 3D reconstruction system acquire and retain useful information to understand the geometry of a scene from partial views under a limited computation budget? Existing active view acquisition methods typically estimate uncertainty over observed or instantiated geometry, limiting their ability to reason about unseen structure, while long-horizon reconstruction methods often retain redundant observations. We introduce Matisse, a training-free framework that unifies active reconstruction and keyframe selection by leveraging evidence provided by a pretrained generative 3D model. Matisse estimates Evidential Uncertainty from cross-attention evidence associated with 3D latent tokens and derives an Evidential Information Gain to guide both view acquisition and keyframe selection based on the expected reduction in posterior entropy. Matisse supports multi-object scenes through occlusion-aware, object-balanced aggregation and propagates uncertainty through intermediate latents to avoid full reconstruction during planning. Matisse reduces Chamfer distance by 12.7%, 3.8%, and 9.2% on GSO30, YCB-V, and Replica, respectively, relative to the best baseline on each dataset, and achieves a $1.50\times$ end-to-end speedup over the best active reconstruction baseline on GSO30 with the same reconstruction backend. In the GSO30 keyframe selection experiment for long-horizon reconstruction, Matisse achieves comparable Chamfer distance using 14% of the input views compared with Stream3D.
Comments20 pages, 7 figures, 7 tables. Project page: https://xihangyu630.github.io/matisse/