arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15798cs.LGstat.ML

语言模型的交叉熵风险估计:不一致性必然是稠密的,留出法也不例外

Cross-Entropy Risk Estimation for Language Models: Inconsistency Must Be Dense, and the Holdout Method Is No Exception

Hanti Lin

首次发表
浏览论文内容

中文总结 AI 辅助

该研究指出语言模型的每词交叉熵风险无法被一致估计,即使采用留出法也存在稠密的不一致状态,提出两种需付出代价的解决途径。

中文摘要 AI 辅助

语言模型通过其留出的每词交叉熵风险进行比较,这是规模定律所拟合的量。我们证明该风险无法被一致估计。一致性(即收敛到被估计量)是相对于“可能的世界状态”定义的:由数据生成分布与我们最终训练的模型组成的对。对模型以及数据生成机制进行量化至关重要,因为决定模型风险是否可估计的是其权重诱导分布的尾部性质,而没有任何样本能揭示这一性质。每词交叉熵风险难以估计源于一个拓扑事实:在所有可能的状态中,有限风险和无限风险彼此任意接近。因此,没有任何估计量——不仅仅是留出平均——能在风险被定义的每个状态上保持一致。更糟的是,即使限制预期序列长度和全支撑模型,不一致估计依然存在;在该受限场景下,发生不一致的状态甚至是稠密的。我们确定了两种可行的解决途径,但均需付出代价。途径1:使用有限上下文窗口,我们可以对模型的下一词概率设置下限,使其风险在数据生成分布具有有限预期序列长度时恰好有限——这为原本出于计算考量的选择提供了新的统计依据,不过其替代的假设本身无法通过任何检验。途径2:仅在风险低于预先设定阈值时报告风险可恢复一致性,且不会对模型选择的实际需求造成损失——但我们需认识到估计目标已被修正。

英文摘要

Language models are compared by their held-out per-token cross-entropy risk---the quantity scaling laws are fitted to. We show that it cannot be consistently estimated. Consistency, or convergence to the estimand, is defined relative to a \emph{possible state of the world}: a pair consisting of a data-generating distribution and a model we turn out to train. Quantifying over models as well as data-generating mechanisms is essential, because what decides whether a model's risk is estimable is a tail property of the distribution its weights induce, which no sample reveals. The per-token cross-entropy risk is hard to estimate because of a topological fact: among the possible states, finite risk and infinite risk each lie arbitrarily close to every instance of the other. Consequently no estimator---not merely the holdout average---is consistent at every state at which the risk is defined. Worse, inconsistent estimation persists under both bounding the expected sequence length and restricting to full-support models; and in that restricted setting the states at which inconsistency occurs are even dense. Two interesting ways out are identified, and neither is free. Way out 1: using a bounded context window, we can floor a model's next-token probabilities, making its risk finite exactly when the data-generating distribution has finite expected sequence length---a new, statistical rationale for a choice that was made on computational grounds, though the assumption it substitutes is itself beyond the reach of any test. Way out 2: reporting the risk only when it falls below a threshold fixed in advance restores consistency, at no cost to what model selection actually requires---but we need to recognize that the goal of estimation is revised.

发表机构

  • University of California, Davis(加利福尼亚大学戴维斯分校)

机构由 AI 辅助整理,请以论文原文为准。

↑