arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不变测度推理器:用于隐式推理的稳定表示

Invariant-Measure Reasoners: Stable Representations for Latent Reasoning

Yuto Inui, Takuya Konishi, Yoshinobu Kawahara

arXiv 2610.10996首次发表:更新:

发表机构

The University of Osaka; RIKEN(大阪大学; 理化学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对隐式推理模型随循环深度增加出现的预测不稳定问题,提出不变测度推理器(ImR)框架,通过利用隐状态的长期不变测度提升预测稳定性与准确率,在迷宫和数独任务中效果显著。

AI 中文摘要

隐式推理模型会使用同一个循环块反复更新隐状态。随着循环深度增加,隐状态序列可能收敛到状态空间的一个紧子集,而不一定收敛到不动点。现有模型通常通过将预测头应用于单个隐状态来进行预测。然而,即使经过多次更新,隐状态仍可能持续变化,这可能导致不同循环深度下的预测不稳定。为解决这种不稳定性,我们提出不变测度推理器(ImR),这是一种使用不变测度作为稳定表示的框架。该测度描述了隐状态在紧子集上的长期分布,且在循环块的更新下保持不变。ImR通过预测头在该测度下的期望进行预测。我们以两种方式使用ImR:仅对现有模型的预测头进行微调,以及从头开始训练模型。两种方法均降低了预测不稳定性,并在迷宫和数独任务的多种设置中提高了准确性。在部分设置中,使用ImR训练的模型比现有模型更频繁地表现出非不动点行为,但即便如此仍能达到高准确率,这与现有模型不同。这些结果表明,ImR能够利用原本会导致不稳定的动力学来进行隐式推理。

英文摘要

Latent reasoning models repeatedly update a latent state using the same recurrent block. As the recurrent depth increases, the sequence of latent states may converge to a compact subset of the state space without necessarily converging to a fixed point. Existing models typically predict by applying a prediction head to a single latent state. However, the latent state can continue to change even after many updates, potentially making predictions unstable across recurrent depths. To address this instability, we introduce invariant-measure reasoners (ImR), a framework that uses an invariant measure as a stable representation. This measure describes the long-run distribution of latent states on the compact subset and is invariant under updates by the recurrent block. ImR predicts from the expectation of the prediction head's output under this measure. We use ImR in two ways: fine-tuning only the prediction head of existing models and training models from scratch. Both approaches reduce prediction instability and improve accuracy in many settings on maze and Sudoku tasks. In some settings, models trained with ImR exhibit non-fixed-point behavior more frequently than existing models yet achieve high accuracy even with such behavior, unlike existing models. These results suggest that ImR can leverage otherwise destabilizing dynamics for latent reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑