arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MVDG:在无约束真实世界图像上的高效多视图三维消歧

MVDG: Efficient Multi-view 3D Disambiguation on Unconstrained Real-World Images

Hanyuan Xiao, Gonglin Chen, Haolin Xiong, Wenbin Teng, Haiwei Chen, Yajie Zhao

arXiv 2610.01098首次发表:更新:

发表机构

University of Southern California; Institute for Creative Technologies(南加州大学; 创意技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大规模野外3D重建中的“替身”幻觉匹配问题,提出基于VGGT的多视图消歧框架MVDG,通过3D感知特征单次推理减少成对比较,利用伪成对训练稳定微调,提升SfM精度与速度。

AI 中文摘要

不同但视觉上相似的3D表面之间的幻觉匹配——即“替身”(doppelgangers)——仍然是大规模、野外3D重建和视觉定位的基本障碍。先前的工作通过成对分类器缓解了这一问题,但这种设计限制了多视图上下文推理,并为下游的运动恢复结构(SfM)带来了O(n^2)的推理复杂度。我们提出了MVDG,一个基于3D基础模型VGGT构建的可扩展多视图消歧框架,它能够联合推理任意数量的多视图图像。通过整合3D感知的多视图特征,我们的方法通过单次编码和解码视图来减少对成对比较的依赖。我们进一步观察到,在噪声监督下直接对VGGT进行多视图微调可能不稳定;受Doppelgangers中标签模糊性的启发,我们从AerialMegaDepth构建了一个伪成对训练集,并表明在采样子集上进行微调能产生稳定的优化和对保留场景的强泛化能力。最后,由于完整的SfM评估(即使使用更快的流水线如GLOMAP)仍然昂贵,我们处理一个伪成对数据集以进行高效验证;我们推导出常规SfM指标与该伪成对测试上分类准确率之间的预测关系。实验表明,我们的方法在实现可比较的成对准确率的同时,在SfM准确性和推理速度上均优于基线。

英文摘要

Illusory matches between distinct yet visually similar 3D surfaces--doppelgangers--remain a fundamental obstacle for large-scale, in-the-wild 3D reconstruction and visual localization. Prior work mitigates this issue with pairwise classifiers, but this design limits multi-view contextual reasoning and incurs O(n^2) inference complexity for downstream structure-from-motion (SfM). We present MVDG, a scalable multi-view disambiguation framework built on the 3D foundation model VGGT, which jointly reasons over an arbitrary number of multiview images. By incorporating 3D-aware multi-view features, our method reduces dependence on pairwise comparisons by encoding and decoding views in a single pass. We further observe that direct multi-view fine-tuning of VGGT can be unstable under noisy supervision; motivated by label ambiguity in Doppelgangers, we construct a pseudo-pairwise training set from AerialMegaDepth and show that fine-tuning on sampled subsets yields stable optimization and strong generalization to held-out scenes. Finally, because full SfM evaluation (even with faster pipelines such as GLOMAP) remains expensive, we process a pseudo-pairwise dataset for efficient validation; we derive a predictive relationship between regular SfM metrics and the classification accuracy on this pseudo-pairwise test. Experiments show that our method achieves comparable pairwise accuracy while improving both SfM accuracy and inference speed over baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑