arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

潜在递归LLM系统的原则性思考

Principled Thoughts for Latent Recursive LLM Systems

Fahd Seddik, Fatemeh Fard

arXiv 2609.36159首次发表:更新:

AI 中文总结

针对潜在递归LLM系统,提出REST训练目标,将有效思考的四个属性转化为可微损失,提升准确率与收敛性,增强潜在通信可解释性。

AI 中文摘要

大型语言模型可以通过在自身隐藏状态上递归或在智能体之间传递这些状态,在连续空间中而非解码文本中进行推理,而训练仅监督最终解码答案的交叉熵(CE),不约束思考过程。理论和实证分析确立并证实了仅CE训练的四个失败模式,这些失败导致正确答案概率降低,例如不同问题间思考的坍缩和保留无关信息。我们引入REST(表示监督思考),一种训练目标,将有效思考表示的四个属性(因果性、最小性、可分离性和稳定性)转化为添加到CE的可微损失。我们在潜在单智能体和多智能体系统中实例化它,无需架构更改或在推理时增加参数。在涵盖数学、科学、医学和代码生成的7个基准上,使用相同的训练数据、计算量和潜在预算,REST在智能体设置和模型规模上相比仅CE训练将准确率提升高达7.5个百分点,并将最终答案的收敛性提升30%。此外,REST思考编码了更多实现正确答案所需的信息,解码它们能更好地恢复智能体的预期输出,这使得潜在通信更易于解释。项目网站:此https URL

英文摘要

Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. Theoretical and empirical analyses establish and confirm four failures of CE-only training that lead to a lower probability of the correct answer such as collapsing thoughts across distinct questions and retaining irrelevant information. We introduce REST (REpresentation-Supervised Thoughts), a training objective that turns four properties of a valid thought representation (causality, minimality, separability, and stability) into differentiable losses added to CE. We instantiate it in latent single-agent and multi-agent systems, without architectural changes or added parameters at inference. Across 7 benchmarks spanning mathematics, science, medicine, and code generation, with the same training data, compute, and latent budget, REST increases accuracy over CE-only training across agent settings and model sizes by up to 7.5 percentage points and convergence on a final answer by 30\%. Furthermore, REST thoughts encode more of what is required to achieve the correct answer, and decoding them better recovers the intended output of the agent, which makes latent communication easier to interpret. Project Website: https://fard-lab.github.io/REST

CommentsProject website: https://fard-lab.github.io/REST

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑