MuViSeg:基于密集几何先验的多视图段对应
MuViSeg: Multi-View Segment Correspondences from Dense Geometry Priors
浏览论文内容
中文总结 AI 辅助
研究针对经典图像对应与对象级推理的差距,基于段级匹配范式提出三个学习匹配头,包括LightGlue风格注意力头、DPT风格多尺度融合模块及多视图扩展,在多个数据集和任务中取得显著性能提升。
中文摘要 AI 辅助
经典图像对应在稀疏关键点或密集像素级别解决,但处理这些匹配的系统(如对象级映射、拓扑导航、场景图维护)是对整个对象进行推理。近期工作通过在实例段级别直接匹配来缩小差距:一个类别不可知的分割器对每个图像进行分割,并通过在掩码上池化来自大型3D基础模型的特征来获得每个段的描述符。我们在此段级匹配范式基础上提出三个学习匹配头:一个在冻结的MASt3R描述符上具有双Softmax评分的LightGlue风格注意力头;一个在池化前从VGGT基础模型中暴露分层空间细节的DPT风格多尺度融合模块;以及作为主要贡献的多视图扩展,它同时对从多个视图中提取的段执行联合自注意力,恢复严格成对匹配器无法达到的传递对应。在Replica和Virtual KITTI 2上的分层零样本协议下,LightGlue风格头在相同MASt3R主干上比无参数的Sinkhorn匹配器在Replica上提高了4.85 AUPRC,在Virtual KITTI 2上提高了25.9 AUPRC。在Habitat-Matterport 3D(HM3D)实例图像导航基准测试的RoboHop拓扑导航管道中,无需重新训练,我们多视图变体将成功率从50%提高到70%,LightGlue风格头将SPL从45.7提高到59.1。
英文摘要
Object-level mapping and topological navigation reason about whole objects, yet the correspondences they rely on are computed over keypoints or pixels and only aggregated into objects afterwards. A recent line of work removes this detour by matching directly at the level of instance segments: a class-agnostic segmenter partitions each image, and per-segment descriptors are pooled from large 3D foundation models over the masks. This shifts the open question to how frozen foundation features should be processed once the unit of matching is a segment. We introduce MuViSeg, three learned matching heads: a LightGlue-style attention head on frozen MASt3R descriptors; a DPT-style multi-scale fusion module that exposes layered spatial detail from VGGT before pooling; and, as our main contribution, a joint multi-view head that attends over segments from several views at once, recovering transitive correspondences that pairwise matchers cannot express. In zero-shot evaluation on Replica and Virtual KITTI 2 across 0--180 degree viewpoint changes, our heads consistently improve over a parameter-free Sinkhorn matcher on the same backbone. Across 102 HM3D navigation episodes, direct segment matchers and sparse keypoint aggregation are statistically indistinguishable, while dense matches voted into masks lose 13.7--18.0 SPL.
发表机构
- Applied AI Institute(应用人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。