arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Think Wider: 缓解隐式思维链推理中的潜在秩坍缩

Think Wider: Mitigating Latent Rank Collapse in Implicit Chain-of-Thought Reasoning

Yuwen Hao, Menglin Yang

arXiv 2609.07406首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对隐式思维链推理中潜在状态秩坍缩导致推理多样性下降的问题,提出轻量级谱正则化器WIDER,在训练时惩罚共享方向投影以拓宽表示子空间,实验证明其提升基线性能并揭示有效秩提高等机制。

AI 中文摘要

思维链(Chain-of-Thought, CoT)推理通过引入中间计算步骤来提升大型语言模型的推理能力,但显式的推理过程会增加解码长度、延迟和上下文成本。隐式思维链(Implicit CoT)通过将中间推理过程转移到连续的潜在状态中,提供了一种更高效的替代方案。然而,潜在推理可能不稳定:连续的潜在状态可能变得过于相似,并坍缩至一个共享的主导方向,从而降低推理轨迹的多样性。在本工作中,我们识别了“潜在秩坍缩”(latent rank collapse)现象,并提出了WIDER,一种用于隐式思维链的轻量级谱正则化器。在训练过程中,WIDER估计每条潜在轨迹的共享方向,并惩罚在该方向上的投影,鼓励潜在状态跨越更广泛表示子空间。该方法即插即用,不改变骨干模型、潜在调度和推理时的解码过程。我们进一步将这种坍缩形式化为隐式推理中的几何瓶颈,并将其缓解视为训练时的正则化问题,而非推理时的解码调整。大量实验表明,WIDER优于匹配的隐式思维链基线,而机制分析揭示了更高的有效秩、更低的主导方向能量以及潜在步骤间冗余的减少。这些结果突显了潜在子空间利用作为高效连续推理的重要因素,为分析和改进隐式思维链提供了几何视角。代码可在该https链接获取。

英文摘要

Chain-of-thought (CoT) reasoning improves the reasoning ability of large language models by introducing intermediate computation, but explicit rationales increase decoding length, latency, and context cost. Implicit CoT offers a more efficient alternative by moving intermediate reasoning into continuous latent states. However, latent reasoning can be unstable: successive latent states may become overly similar and collapse toward a shared dominant direction, reducing the diversity of the reasoning trajectory. In this work, we identify $\textit{latent rank collapse}$ and propose $\textbf{WIDER}$, a lightweight spectral regularizer for implicit CoT. During training, WIDER estimates the shared direction of each latent trajectory and penalizes projections onto this direction, encouraging latent states to span a broader representational subspace. The method is plug-and-play and leaves the backbone model, latent schedule, and inference-time decoding procedure unchanged. We further formulate this collapse as a geometric bottleneck in implicit reasoning, casting its mitigation as a training-time regularization problem rather than an inference-time decoding change. Extensive experiments show that WIDER improves matched implicit CoT baselines, while mechanistic analyses reveal higher effective rank, lower dominant-direction energy, and reduced redundancy among latent steps. These results highlight latent subspace utilization as an important factor for efficient continuous reasoning, providing a geometric perspective for analyzing and improving implicit CoT. Code is available at https://github.com/whitesweater/WIDER.

Comments16 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑