arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SpatialCORE:大型视觉-语言模型中的置信度感知接地空间推理

SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models

Rafi Ibn Sultan, Xiangyu Zhou, Md. Sajid Alam Chowdhury, Chengyin Li, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu

arXiv 2609.38716首次发表:更新:

发表机构

Wayne State University; Henry Ford Health; The Ohio State University(韦恩州立大学; 亨利·福特医疗系统; 俄亥俄州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SpatialCORE提出一种后训练框架,利用模型对生成接地的置信度作为学习信号,通过自调节空间奖励和答案门强化自信且准确的接地,提升大型视觉-语言模型的空间推理能力,在多个基准上达到最先进水平。

AI 中文摘要

大型视觉-语言模型(LVLMs)在视觉感知任务上取得了显著进展,但空间推理仍是一个持续的弱点,尤其是对于需要在视觉空间中进行推理的问题。近期的空间推理方法引入了生成的接地(grounding),即模型在推理过程中为任务相关对象预测边界框、掩码或其他定位输出。然而,这些方法通常仅优化最终答案的正确性,即使模型并非基于自信定位的任务相关对象进行推理,正确答案也可能获得奖励。我们提出了SpatialCORE(空间自信推理),一个后训练框架,将模型自身对生成接地的置信度转化为空间推理的学习信号。其核心思想是强化既准确又自信的接地,鼓励模型基于自信定位的任务相关对象进行推理。SpatialCORE通过一种自调节的空间奖励实现这一点,该奖励根据每个预测边界框的坐标令牌置信度对其匹配质量进行加权。一个答案门进一步将接地优化与最终答案的正确性联系起来。SpatialCORE在多个基准测试中,在开源和专门的空间推理模型中取得了最先进的结果,并在零样本设置下有效迁移到未见过的数据分布。源代码可在以下网址获取:https://this https URL。

英文摘要

Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes, masks, or other localization outputs for task-relevant objects as part of their reasoning trace. However, these approaches typically optimize final-answer correctness alone, allowing correct answers to be rewarded even when the model does not reason from confidently localized task-relevant objects. We introduce SpatialCORE (Spatially COnfident REasoning), a post-training framework that turns the model's own confidence in generated grounding into a learning signal for spatial reasoning. Its central idea is to reinforce grounding that is both accurate and confident, encouraging the model to reason from confidently localized task-relevant objects. SpatialCORE realizes this through a self-regulating spatial reward that weights each predicted bounding box's matching quality by its coordinate-token confidence. An answer gate further ties grounding optimization to final-answer correctness. SpatialCORE achieves state-of-the-art results among open-source and specialized spatial reasoning models across diverse benchmarks, and transfers effectively in zero-shot settings to unseen data distributions. The source code is available at https://github.com/rafiibnsultan/SpatialCORE.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑