arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

渲染以推理:新视角语义预测提升VLM空间理解

Render to Reason: Novel-View Semantic Prediction Improves Spatial Understanding in VLMs

Yuqun Wu, Yao Xiao, Chuhang Zou, Shenlong Wang, Derek Hoiem

arXiv 2610.05417首次发表:更新:

发表机构

University of Illinois at Urbana-Champaign; Meta(伊利诺伊大学厄巴纳-尚佩恩分校; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对几何特征融合训练信号不足的问题,提出新视角语义渲染辅助任务,联合几何与视觉路径,在三个基准上显著提升VLM空间理解,超越开源方法。

AI 中文摘要

近期工作通过预训练3D模型为视觉语言模型(VLM)注入几何特征,期望几何信号能增强空间推理能力。然而,我们发现简单融合几何特征并在标准空间问答(QA)上训练,仅在高层次多跳任务上带来微小改进。我们将这一差距归因于训练信号问题:标准空间QA在很大程度上可通过视觉特征和语言先验来回答,因此几何路径获得的梯度较弱,无法与视觉特征有效整合。为提供需要几何信息的训练信号,我们提出\textbf{新视角语义渲染}作为辅助训练任务,要求模型预测未观测视角的语义布局,灵感来源于人类在空间推理中通过心理模拟新视角的能力。该任务鼓励两条路径的联合使用:几何提供姿态相关的可见性,视觉提供语义内容。我们的辅助任务在三个基准上均较几何增强基线带来一致改进(VSI-Bench上提升最高+1.6,ReVSI上+2.2,我们提出的3D-Point-QA数据集上+2.9),且完整模型在VSI-Bench和ReVSI上超越了先前开源方法。项目页面:此https URL。

英文摘要

Recent works augment Vision-Language Models with geometry features from pretrained 3D models, expecting that the geometric signal will boost spatial reasoning. However, we find that simply fusing geometry features and training on standard spatial QA yields only marginal improvements on high-level multi-hop tasks. We attribute this gap to a training-signal problem: standard spatial QA can be largely answered from visual features and language priors, so the geometry pathway receives weak gradients and fails to integrate with the visual features. To provide a training signal that requires geometry, we propose \textbf{novel-view semantic rendering} as an auxiliary training task that requires the model to predict the semantic layout of an unobserved viewpoint, inspired by humans' ability to mentally simulate novel viewpoints during spatial reasoning. This task encourages joint use of both pathways: geometry provides pose-dependent visibility, while vision provides semantic content. Our auxiliary task yields consistent improvements over the geometry-augmented baseline across all three benchmarks (up to +1.6 on VSI-Bench, +2.2 on ReVSI, +2.9 on our 3D-Point-QA dataset) and our full model surpasses prior open-source methods on VSI-Bench and on ReVSI. Project page: https://yuqunw.github.io/Render2Reason/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑