发表机构
NYU Shanghai Leonard N. Stern School of Business, New York University(纽约大学上海纽约大学斯特恩商学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究语言模型中信息分配定律,通过单行中心化恒等式等方法,揭示宏观类别结构与令牌内上下文方差占比,理论上说明权重衰减影响,证明二元类别下限,表明定律在多方面成立及预训练中类别份额变化规律。
AI 中文摘要
语言模型表示中的类内方差通常被视为不完全的神经坍缩。我们认为它是分配的信息存储,且这种分配遵循一条定律。一条单行的中心化恒等式使一类单纯形等角紧框架断言无效,包括我们自己之前的断言。在14个模型的无量纲方差份额中,宏观类别结构仅承载4 - 12%的表示方差,令牌内上下文承载79 - 91%,在100倍参数范围内稳定。理论上,令牌级权重衰减按类型计数而非出现质量惩罚类别,将下一个令牌预测简化为不平衡的K类问题,其最优值按类型计数对类别范数排序。一个针对二元类别的相反下限迫使类内离散至少与条件互信息I(令牌;上下文|类别)成比例。该定律成立:身份离散而非总方差在每个测试模型和划分中跟踪此信息,在无模型估计下甚至跨模型,其中一个模型的信息预测另一个模型的离散;并且在预训练过程中,类别份额会超调、衰减并部分恢复,因为它必须承载的信息从未离开。
英文摘要
Within-class variance in language-model representations is commonly read as incomplete neural collapse. We argue it is allocated information storage, and that the allocation obeys a law. A one-line centering identity voids a family of simplex equiangular-tight-frame claims, including our own earlier ones; in dimensionless variance shares across 14 models, macro-category structure carries only 4-12% of representational variance and within-token context carries 79-91%, stable across a 100x parameter range. On the theory side, token-level weight decay penalizes a category in proportion to its type count, not its occurrence mass, reducing next-token prediction to an imbalanced K-class problem whose optimum orders category norms by type count. A converse floor, proved for binary categories, forces within-category dispersion to be at least proportional to the conditional mutual information I(token; context | category). The law holds: identity dispersion, not total variance, tracks this information across every tested model and partition, under a model-free estimate and even across models, where one model's information predicts another's dispersion; and over pretraining the category share overshoots, decays, and partially recovers, because the information it must carry never left.