arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30072stat.MLcs.LGmath.STstat.TH

通过Token预测学习表示:几何、近似与下游保证

Learning Representations through Token Prediction: Geometry, Approximation, and Downstream Guarantees

Shulei Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究构建关联Token预测与表示几何等的统计框架,证明其能组织Token嵌入、优化上下文表示,并为下游任务提供保证,解释了Token预测可生成有用表示的原理。

中文摘要 AI 辅助

Token预测是现代语言模型的核心预训练目标,尽管其在实践中取得了成功,但为何Token预测能学习到广泛有用的表示仍未被完全理解。我们开发了一个将Token预测与表示几何、编码器近似及下游性能关联起来的统计框架。在softmax预测头下,我们证明准确的Token预测会根据不同Token类型出现的上下文分布之间的相似性(由Hellinger距离衡量)来组织Token嵌入,其明确误差由预测准确率和Token频率决定。同时,上下文表示为目标Token相对于这些嵌入的条件分布提供了低维坐标。我们进一步引入自一致性原理,表明共享表示块的重复应用可逐步优化上下文表示,且不会引入额外的块参数。在具有相同预测准确率的表示中,这种循环构造倾向于那些可从其上下文稳定重构的表示。最后,我们为Token生成、Token社区恢复及线性探测分类建立了下游保证,展示了预测准确率和恢复的几何如何转化为超出预训练目标的性能。这些结果共同解释了简单的Token预测目标如何能恢复语义几何并产生广泛有用的表示,一项受控模拟说明了这些理论机制。

英文摘要

Token prediction is a central pre-training objective for modern language models. Despite its empirical success, why token prediction learns broadly useful representations remains incompletely understood. We develop a statistical framework connecting token prediction with representation geometry, encoder approximation, and downstream performance. Under a softmax prediction head, we show that accurate token prediction organizes token embeddings according to similarities between the distributions of contexts in which different token types appear, as measured by Hellinger distance, with explicit errors governed by prediction accuracy and token frequency. Meanwhile, the contextual representation provides a low-dimensional coordinate for the conditional distribution of the target token relative to these embeddings. We further introduce a self-consistency principle showing that repeated applications of a shared representation block can progressively refine the contextual representation without introducing additional block parameters. Among representations with the same prediction accuracy, this recurrent construction favors those that can be stably reconstructed from their contexts. Finally, we establish downstream guarantees for token generation, token community recovery, and classification by a linear probe, showing how prediction accuracy and recovered geometry translate into performance beyond the pre-training objective. Together, these results explain how the simple objective of predicting tokens can recover semantic geometry and produce broadly useful representations. A controlled simulation illustrates the theoretical mechanisms.

发表机构

  • University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

↑