arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34731stat.MLcs.ITcs.LGmath.IT

马尔可夫数据下下一词预测的信息论分析

Information-Theoretic Analysis of Next-Token Prediction under Markovian Data

Masoud Kavian, Abdellatif Zaidi, Milad Sefidgaran

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出信息论框架分析时间相关数据下下一词预测的泛化界,揭示上下文长度、模型复杂度与时序混合的作用,并通过实验表明长上下文扩大泛化差距且存在约24小时的有效记忆尺度。

中文摘要 AI 辅助

我们为时间相关数据下的下一词预测中的泛化问题开发了一个信息论框架。我们考虑由有限记忆马尔可夫过程生成的独立轨迹,并区分由互信息量化的算法依赖性与由混合特性刻画的时序依赖性。对于交叉熵损失,我们利用Donsker--Varadhan变分表示和针对马尔可夫链的McDiarmid型集中不等式推导了一个期望泛化界。一个改进版本通过历史状态过程的混合特性捕捉了上下文长度与时序混合的联合效应。随后,我们通过率失真表述扩展该界,用表示学习模型在泛化差距中规定失真内所需的最小信息率替代互信息,从而为连续假设空间上的确定性算法提供有信息量的保证。对于基于间隔的预测,我们通过带噪低维压缩推导了线性和自注意力下一词预测器的显式界,揭示了上下文长度、模型复杂度、样本量、间隔和时序混合的作用。在TinyStories和ETTh2上的实验表明,更长的上下文可以同时降低训练和测试损失,但通常更多地降低训练损失,从而扩大了泛化差距。一项补充性的ETTh2分析确定了约24小时的有效预测记忆尺度,超过该尺度后没有统计上支持的性能提升,这为更大上下文下测试性能饱和提供了合理解释。

英文摘要

We develop an information-theoretic framework for generalization in next-token prediction under temporally dependent data. We consider independent trajectories generated by finite-memory Markov processes and distinguish algorithmic dependence, quantified by mutual information, from temporal dependence, characterized by mixing. For cross-entropy loss, we derive an expected generalization bound using the Donsker--Varadhan variational representation and a McDiarmid-type concentration inequality for Markov chains. A refinement captures the joint effect of context length and temporal mixing through the mixing properties of the history-state process. We then extend the bound through a rate--distortion formulation, replacing mutual information with the minimum information rate required to represent the learned model within a prescribed distortion in the generalization gap, yielding informative guarantees for deterministic algorithms over continuous hypothesis spaces. For margin-based prediction, we derive explicit bounds for linear and self-attention next-token predictors via noisy low-dimensional compression, revealing the roles of context length, model complexity, sample size, margin, and temporal mixing. Experiments on TinyStories and ETTh2 show that longer contexts can reduce both training and test losses, but typically reduce training loss more, enlarging the generalization gap. A complementary ETTh2 analysis identifies an effective predictive-memory scale near 24 hours, with no statistically supported improvement beyond this scale, offering a plausible explanation for test-performance saturation at larger contexts.

发表机构

  • Huawei Paris Research Center(华为巴黎研究中心)
  • Université Gustave Eiffel(古斯塔夫·埃菲尔大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑