arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于多视图空间理解的搜索-视图推理

Seek-and-View Reasoning for Multi-View Spatial Understanding

Qixiang Chen, Cheng Zhang, Fucai Ke, Chi-Wing Fu, Jianfei Cai, Jingwen Ye

arXiv 2610.11810首次发表:更新:

发表机构

Monash University; The Chinese University of Hong Kong(莫纳什大学; 香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有多视图空间推理方法的跨视图对齐脆弱等问题,提出无需微调的模型无关框架Vantage,结合VLM与3D基础模型,经实验在五个基准上提升了多视图空间理解性能。

AI 中文摘要

现有的多视图空间推理方法大多基于稀疏输入视图运行,因此视觉语言模型(VLMs)被限制在这些固定视图内理解场景并推断空间关系,导致跨视图对齐脆弱以及几何到语言的瓶颈。为解决这些问题,我们提出了一种新颖的搜索-视图推理方法,通过定位与问题相关的视图来寻找隐式跨视图空间证据,以支持空间推理。为实现该方法,我们提出了Vantage,这是一种无需微调、与模型无关的推理框架,它将VLM与3D基础模型配对:首先是基于视点的推理阶段,用于问题分析和视图规划;随后是基于几何的证据增强阶段,以有效合成视觉证据并将其融入最终推理。在六个VLMs上进行的综合实验表明,该方法在五个基准上取得了一致的改进,且无需微调。总体而言,通过基于视图的推理揭示空间证据,Vantage可大幅减少对基于语言的跨视图对齐的依赖,提升多视图空间理解能力。我们的代码可在该https URL获取。

英文摘要

Existing approaches to multi-view spatial reasoning operate largely on sparse input views. Vision-language models (VLMs) are thus restricted to understand a scene and infer spatial relations within these fixed views, leading to fragile cross-view alignment and geometry-to-language bottleneck. To address these issues, we formulate a novel Seek-and-View reasoning approach to find implicit cross-view spatial evidence by locating a question-relevant view to support the spatial reasoning. To realize this approach, we propose Vantage, a training-free model-agnostic reasoning framework that pairs a VLM with a 3D foundation model: a viewpoint-grounded reasoning stage for question analysis and view planning, followed by a geometry-grounded evidence augmentation stage to effectively synthesize and incorporate visual evidence into the final reasoning. Comprehensive experiments on six VLMs demonstrate consistent improvements on five benchmarks without fine-tuning. Overall, by revealing spatial evidence through view-grounded reasoning, Vantage can largely reduce reliance on language-based cross-view alignment and improve multi-view spatial understanding. Our code is available at https://github.com/q1xiangchen/Vantage.

CommentsProject page: https://seekandview2026.github.io; Code: https://github.com/q1xiangchen/Vantage

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑