一个向量能容纳多少思想?叠加推理的容量
How Should Reasoning Be Organized in a Transformer's Latent Space?
浏览论文内容
中文总结 AI 辅助
本文研究连续潜在状态中推理叠加的容量问题,证明保留完整历史的累积叠加在固定维度下优于仅保留当前前沿的叠加,并提出均匀累积加权作为极小极大最优的稳健记忆策略。
中文摘要 AI 辅助
大型语言模型通过多步推理中的中间计算来解决难题。传统的思维链将这些计算编码为词元。最近的连续和循环方法则将部分计算移入固定维度的潜在状态中,其中单个思想可以叠加多个备选方案。这引发了一个基本的设计问题:随着推理的进行,连续思想应保留什么?一种直观的方法是丢弃过去的计算,仅保留当前的推理前沿。存储更多项目似乎会稀释状态并浪费有限的表现容量。我们表明这种直觉可能是错误的。在相同的下游计算下,保留完整推理历史的累积叠加可能比仅保留当前备选方案的前沿叠加需要更低的表现维度。在固定的隐藏宽度下,这一优势使潜在推理器能够保留更多有效证据,区分更多合理的下游结果,并延迟压缩状态变得不可靠的临界点。这种反直觉效应的出现是因为信息丰富的历史成分相互连贯地增强,而不相关的备选方案则带来随机干扰。这一视角也回答了一个实际的设计问题:当模型对潜在状态内积累的记忆的未来用途未知时,应如何对其加权?在可复用的加权叠加中,优先考虑少量近期或显著项目会导致弱表示的记忆,从而成为后续注意力的瓶颈。均匀累积加权避免了这一缺陷,我们证明其对于稳健的未来推理是极小极大最优的。我们的结果将叠加从一种观察到的潜在空间效应转变为设计原则:平衡的累积记忆使固定的表现预算能够支持更可靠、可复用的计算。
英文摘要
Continuous reasoning has emerged as a promising way to improve reasoning in large language models (LLMs). Yet we still lack a clear principle for deciding what a latent state should preserve. Reasoning by superposition shows that a single latent state can encode several search alternatives and expand them in parallel. We ask how those states should be weighted as reasoning proceeds. A natural choice is to preserve only the states active at the frontier step, since keeping every reached state appears to spread a limited hidden width too thin. We show that the opposite can hold. When later computation draws on several reached states, a cumulative state can guide attention correctly at a smaller hidden width than a frontier state that stores fewer states. At the same width, the cumulative state therefore keeps more intermediate states available for later reasoning. More generally, equal cumulative weights are optimal when future queries are unknown and remain close to the best task-specific weights when those queries are known. Experiments with two-layer and GPT-2 Transformers reproduce the predicted width advantage and show that unequal weights fail first on the states that receive the least weight. This suggests a important principle: keep reached states equally weighted, and restore equal weights as computation proceeds.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- Zhongguancun Academy and Technology of China(中关村科技学院)
机构由 AI 辅助整理,请以论文原文为准。