arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29733cs.CV

XDG:加速视觉消歧

XDG: Accelerated Visual Disambiguation

Gonglin Chen, Ben Southall, Hanyuan Xiao, Wenbin Teng, Haolin Xiong, Tianwen Fu, Junyi Ouyang, Kshitij Singh Minhas, Supun Samarasekera, Rakesh Kumar, Yajie Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

XDG是适配可扩展SfM的高效视觉消歧模型,通过微调Depth Anything 3并复用其相机令牌实现配对分类,在保持精度的同时将推理速度提升超3倍,大幅缩短大规模场景消歧耗时。

中文摘要 AI 辅助

视觉混叠,又称孪生问题,仍是运动恢复结构(SfM)的核心挑战:视觉相似但物理不同的表面会产生错误的图像匹配,降低重建质量。现有研究通过几何感知的基础模型特征缓解该问题,但在骨干网络上方部署了重型Transformer分类器,导致大规模消歧成本高昂。本文提出XDG,一种为可扩展SfM设计的高效视觉消歧模型。核心观察是3D基础模型已具备视觉消歧所需的跨视图几何推理能力,因此孪生分类应直接适配骨干网络表示,而非在单独的重型解码器中重新学习配对推理。XDG通过轻量LoRA适配器微调Depth Anything 3,并将其相机令牌重新用作紧凑的配对级分类令牌;一个紧凑的MLP头预测候选图像对是否观测到同一3D表面。大量实验表明,XDG在精度与效率间取得良好平衡:在配对和重建基准上与最先进的消歧方法表现相当,且推理速度提升超3倍;在包含数千张图像的单个LaMAR场景中,XDG节省超10小时的视觉消歧处理时间。代码可在该URL获取。

英文摘要

Visual aliasing, also known as the doppelganger problem, remains a key challenge for structure-from-motion (SfM): visually similar but physically distinct surfaces can produce incorrect image matches and degrade reconstruction quality. Previous work mitigates this issue with geometry-aware foundation-model features, but places a heavy transformer classifier on top of the backbone, making large-scale disambiguation expensive. We introduce XDG, an efficient visual disambiguation model designed for scalable SfM. Our key observation is that a 3D foundation model already performs the cross-view geometric reasoning necessary for visual disambiguation, so doppelganger classification should adapt the backbone representation directly rather than relearn pair reasoning in a separate heavy decoder. XDG fine-tunes Depth Anything 3 with lightweight LoRA adapters and repurposes its camera tokens as compact pair-level classification tokens. A compact MLP head predicts whether a candidate image pair observes the same 3D surface. Extensive experiments show that XDG provides a favorable accuracy-efficiency tradeoff: it remains competitive with the state-of-the-art disambiguation method across pairwise and reconstruction benchmarks and delivers more than a 3x inference speedup. On individual LaMAR scenes containing thousands of images, XDG saves more than 10 hours of visual disambiguation processing. Code is available at https://github.com/xtcpete/xdg.

发表机构

  • USC Institute for Creative Technologies(南加州大学创意技术研究所)
  • University of Southern California(南加州大学)
  • SRI International(SRI国际公司)

机构由 AI 辅助整理,请以论文原文为准。

↑