发表机构
Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过Anthropic的叠加玩具模型,揭示了模型宽度和数据统计如何共同决定特征表征的相变与损失缩放,为理解大型模型中的表征扩展提供了理论依据。
AI 中文摘要
大型语言模型被认为通过隐藏空间中的向量来表示特征,该空间的维度由模型的宽度决定。叠加,即通过让表征向量重叠来表示比宽度更多的特征,是表征向量如何组织的主要解释。然而,当特征数量和宽度都很大时,模型宽度和数据统计如何决定表征向量的配置以及由此产生的损失,仍不太清楚。在这里,我们展示了在Anthropic的叠加玩具模型中,增加宽度驱动了一个连续相变,从部分表征相(其中只有一部分特征获得可观的表征向量,其余特征消失)到完全表征相(其中每个特征都被表征)。我们通过部分随机投影近似理论预测,并且实验证实,临界宽度随活跃特征数量线性增长,直至一个对数因子。损失缩放随相变而变化:在临界宽度以下,损失随活跃特征数量线性增长,并以数据统计设定的形式弱依赖于宽度;在临界宽度以上,损失随活跃特征数量近似二次增长,并与宽度成反比衰减。非均匀的激发概率延迟了相变并降低了损失,因为更频繁的特征占据更多空间。我们的结果提供了模型宽度和数据统计如何共同塑造表征和损失的说明,这是理解大型模型中表征缩放的一步。
英文摘要
Large language models are thought to represent features by vectors in a hidden space of dimension given by the model's width. Superposition, in which more features are represented than the width by letting representation vectors overlap, is a leading account of how representation vectors are organized. However, how model width and data statistics determine the configuration of representation vectors and the resulting loss when the number of features and the width are large remains less understood. Here we show, in Anthropic's toy model of superposition, that increasing the width drives a continuous phase transition from a partial-representation phase, where only a subset of features receives appreciable representation vectors while the rest vanish, to a full-representation phase, where every feature is represented. Our theory via a partial random projection approximation predicts, and experiments confirm, that the critical width grows linearly with the number of active features up to a logarithmic factor. The loss scaling changes across the transition: below the critical width, the loss grows linearly with the number of active features and depends weakly on the width in a form set by data statistics; above it, the loss grows approximately quadratically with the number of active features and decays inversely with the width. Non-uniform firing probabilities delay the transition and lower the loss, as more frequent features occupy more space. Our results provide an account of how model width and data statistics jointly shape representations and loss, a step toward understanding representation scaling in large models.
Comments35 pages, 25 figures