arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06746cs.AIcs.CLcs.CVcs.LG

通过潜在空间进行推理!使潜在视觉推理成为必要

Reason Through the Latent! Making Latent Visual Reasoning Necessary

Suhyeong Park, Junha Jung, Jaewoo Kang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出因果视觉循环推理(CVRR),通过强制循环计算作为图像条件路径,确保潜在视觉推理真正依赖视觉信息,并在多个基准上验证其有效性。

中文摘要 AI 辅助

潜在视觉推理旨在通过隐藏状态计算而非显式文本思维链来执行多模态推理。然而,视觉信息存在于潜在状态中并不意味着模型在生成答案时实际上依赖该状态,尤其是在存在其他图像条件路径可用的情况下。我们引入了因果视觉循环推理(CVRR),它在保留预训练视觉能力的同时,使循环计算成为预测所必需的图像条件路径。CVRR在预训练视觉语言模型整合图像后,从问题隐藏状态初始化循环,然后反复更新该状态,同时重新读取相同的固定视觉证据。在解码之前,移除视觉状态和原始多模态KV缓存,使得只有最终的循环状态携带图像条件信息到答案。在V*、MMVP、BLINK和MME-RealWorld-Lite基准上,CVRR在此严格接口下保持强性能,而兼容的潜在推理器即使在相同约束下重新训练也无法恢复可比的视觉能力。因果干预进一步表明,当问题保持不变时,预测对循环内容保持敏感,并且持续的视觉证据因果性地修正循环轨迹。这些结果区分了潜在信息性与实际用于预测的潜在计算。

英文摘要

Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textual chains of thought. However, visual information being present in a latent state does not imply that the model actually relies on that state when producing its answer, especially when alternative image-conditioned paths remain available. We introduce Causal Visual Recurrent Reasoning (CVRR), which preserves pretrained visual competence while making recurrent computation the required image-conditioned path to prediction. CVRR initializes recurrence from the question hidden state after the pretrained vision-language model has incorporated the image, then repeatedly updates this state while re-reading the same fixed visual evidence. Before decoding, visual states and the original multimodal KV cache are removed so that only the final recurrent state carries image-conditioned information to the answer. Across the $V^*$, MMVP, BLINK, and MME-RealWorld-Lite benchmarks, CVRR retains strong performance under this strict interface, while compatible latent reasoners fail to recover comparable visual competence even when retrained under the same constraint. Causal interventions further show that predictions remain sensitive to recurrent content when the question is held fixed, and that persistent visual evidence causally revises the recurrent trajectory. These results distinguish latent informativeness from latent computation that is actually used for prediction.

发表机构

  • Korea University(高丽大学)
  • AIGEN Sciences(AIGEN 科学公司)

机构由 AI 辅助整理,请以论文原文为准。

↑