arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从感知到整合:重新审视视觉-语言模型推理的内部动态

From Perception to Integration: Revisiting the Internal Dynamics of Reasoning in Vision-Language Models

Rong Yu Xu, Prayag Tiwari, Shaolei Zhang

arXiv 2609.34809首次发表:更新:

发表机构

Shenzhen College of International Education; Halmstad University; Renmin University of China(深圳国际交流学院; 哈尔姆斯塔德大学; 中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过受控任务揭示视觉-语言模型在组合推理中的内部动态,提出基于隐藏状态就绪性检测的提前停止策略,在MMStar和RealWorldQA上分别减少79.1%和74.5%的推理令牌,同时提升准确率。

AI 中文摘要

视觉-语言模型(VLMs)能够回答简单的视觉问题,但当一个问题需要多项视觉判断时,它们往往表现困难。我们通过受控任务研究这一差距,这些任务涉及特征绑定、数量感知、空间关系以及模态补全,并设计了一个组合任务(Composite task)将上述任务结合起来。匹配的反事实图像对隔离了回答问题所需的视觉证据变化。在四个模型中,直接答案、隐藏状态读取以及状态干预表明,个体判断可以在无需显式推理的情况下完成,并且对相应状态进行干预可以影响最终答案。在推理过程中,组合答案从隐藏状态中变得可解码,并可从缩短的轨迹中使用,这往往发生在模型自行停止之前。我们训练了一个小型检测器来预测这种就绪状态,并在该点停止推理。在MMStar和RealWorldQA上,这种方法将平均推理令牌数分别减少了79.1%和74.5%,而平均准确率分别提升了3.13和3.30个百分点。这些发现将答案就绪性的内部发展与分配推理计算的实际规则联系起来。

英文摘要

Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterfactual image pairs isolate changes in the visual evidence needed to answer. Across four models, direct answers, hidden-state readouts, and state interventions show that the individual judgments can be made without explicit reasoning and that intervening on the corresponding states can affect the answer. During reasoning, the Composite answer becomes decodable from hidden states and usable from shortened traces, often before the model stops on its own. We train a small detector to predict this readiness and stop reasoning at that point. On MMStar and RealWorldQA, this reduces mean reasoning tokens by 79.1% and 74.5%, while average accuracy rises by 3.13 and 3.30 percentage points, respectively. These findings connect the internal development of answer readiness to a practical rule for allocating reasoning computation.

Comments15 pages, 4 figures. Code: https://github.com/allenxu09/from-perception-to-integration

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑