arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38717cs.CVcs.AI

软空间推理

Soft Spatial Reasoning

Rafi Ibn Sultan, Md. Sajid Alam Chowdhury, Saleh Zare Zade, Chengyin Li, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

针对LVLMs空间推理中的过早离散化问题,提出软空间推理框架,通过AdaptSoft控制器动态调整软度,并利用梯度对齐目标训练,在多个空间基准上超越现有方法。

中文摘要 AI 辅助

大型视觉语言模型(LVLMs)通常通过思维链(CoT)进行空间推理,将中间推理编码为离散语言标记的自回归序列。这种硬性思考要求在每一步都承诺一个单一标记,即使正确的空间解释仍不确定。这种早期承诺构成了过早离散化:不正确的标记选择可能会在后续推理中传播错误。我们提出软空间推理,一种为LVLMs中的空间任务引入软思考的后训练框架。在每个中间推理步骤中,LVLM通过混合标记嵌入而不是选择单一标记来形成连续的软状态,允许多个候选续项影响下一步。然而,适当的软度程度可能因推理步骤而异:保留多个候选可能保留有用的空间解释,但如果这些候选暗示冲突的空间关系,混合它们可能会干扰后续推理。软空间推理的核心是AdaptSoft,一个控制器,它使用当前隐藏状态和预测不确定性来调整每个推理步骤的软度程度。为了训练AdaptSoft,我们引入了一个梯度对齐学习目标,为软度控制提供步骤特定的学习信号,而无需中间推理监督。在多样化的空间基准上,软空间推理优于使用相同骨干的硬性和固定软CoT基线,以及一系列现有的LVLMs。源代码可在https://this https URL获取。

英文摘要

Large Vision-Language Models (LVLMs) commonly perform spatial reasoning through chain-of-thought (CoT), encoding intermediate reasoning as autoregressive sequences of discrete language tokens. Such hard thinking requires committing to a single token at each step, even when the correct spatial interpretation remains uncertain. This early commitment constitutes premature discretization: an incorrect token selection can propagate errors through subsequent reasoning. We propose Soft Spatial Reasoning, a post-training framework that introduces soft thinking for spatial tasks in LVLMs. At each intermediate reasoning step, the LVLM forms a continuous soft state by mixing token embeddings rather than selecting a single token, allowing multiple candidate continuations to influence the next step. The appropriate degree of softness, however, can vary across reasoning steps: retaining multiple candidates may preserve a useful spatial interpretation, but if those candidates imply conflicting spatial relations, mixing them may interfere with subsequent reasoning. At the core of Soft Spatial Reasoning is AdaptSoft, a controller that uses the current hidden state and predictive uncertainty to adapt the degree of softness at each reasoning step. To train AdaptSoft, we introduce a gradient-alignment learning objective that provides a step-specific learning signal for softness control without intermediate reasoning supervision. Across diverse spatial benchmarks, Soft Spatial Reasoning outperforms hard and fixed-soft CoT baselines using the same backbone, as well as a range of existing LVLMs. The source code is available at https://github.com/rafiibnsultan/Soft_Spatial_Reasoning

发表机构

  • Wayne State University(韦恩州立大学)
  • Henry Ford Health(亨利·福特健康)
  • The Ohio State University(俄亥俄州立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑