arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度递归语言模型中的预测动力学

Prediction Dynamics in Depth-Recurrent Language Models

Xinyue Luo, Fei Yu

arXiv 2609.21383首次发表:更新:

发表机构

Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过分解分数变化的几何与组成,解释了深度递归语言模型中中间答案与最终一致但分数变化的现象,并量化了方向与配对对预测深度的影响。

AI 中文摘要

深度递归语言模型通过重复的潜在状态更新来细化预测。为什么中间答案能与最终结果一致,而其分数却持续变化?我们推导出一个尖锐的边际特征,将幅度界限的保守性分解为公共平移、相对于胜者的方向以及每个竞争者更新与其分数差距的配对。在Huginn-3.5B和Ouro-1.4B上,在完整答案文本评分下,考虑更新方向和竞争者配对后,平均最早合格深度比仅去除平移进一步减少了总深度的22.5%-34.4%。这一回顾性比较使用了完整的轨迹。在标签评分下也存在实质性贡献。对于共享的预测分布,我们将公共运动和对比运动正交分离,并通过候选集质量和集内集中度表达公共成分。公共能量和对比能量可以以不同速率衰减,使得增长中的偏好变化份额与缩小的绝对更新共存。这些发现通过观察到的分数变化的几何和组成解释了有限深度下的答案保持。

英文摘要

Depth-recurrent language models refine predictions through repeated latent updates. Why can intermediate answers agree with the endpoint while their scores continue to change? We derive a sharp margin characterization that decomposes the conservatism of a magnitude bound into common translation, direction relative to the winner, and the pairing of each competitor's update with its score gap. Across Huginn-3.5B and Ouro-1.4B, accounting for update direction and competitor pairing reduces the mean earliest qualifying depth by a further 22.5-34.4% of the total depth beyond translation removal under full answer-text scoring. This retrospective comparison uses completed trajectories. Substantial contributions also occur under label scoring. For shared predictive distributions, we separate common and contrast motion orthogonally and express the common component through candidate-set mass and within-set concentration. Common and contrast energies can attenuate at different rates, allowing a growing preference-change share to coexist with shrinking absolute updates. These findings explain finite-depth answer preservation through the geometry and composition of observed score changes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑