arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18989cs.AI

函数存在于方差不存在之处:语言模型计算的任务加权图表

Function Lives Where Variance Doesn't: Task-Weighted Charts of a Language Model's Computation

Alexandre Quemy

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出任务加权图表方法,通过函数度量下的低维坐标拟合,揭示语言模型计算维度随函数而异,且函数与方差分离,在低维压缩时优于方差方法。

中文摘要 AI 辅助

语言模型的计算实际上使用了多少维度?在命名一个函数之前,这个问题是不适定的。任务加权图表使其成为适定的:针对表示的一个选定函数,在其自身的度量下拟合低维坐标系,将蒸馏转化为简单的普通最小二乘法。跨越三个家族的六个模型,参数规模从70M到7B,下一词预测需要残差流宽度的70%到90%才能保持困惑度在完整值的5%以内,这一宽度被语言的稀有尾部消耗,而方差分布无法预测这一点:两个方向承载了GPT-2激活方差的90%,却几乎不承载其功能。维度是随函数而异的:模型自身的不确定性从六个坐标中读取,而完整预测分布需要数百个坐标;并且维度随深度增加。这种分离是可利用的:当只能保留少数维度时,在函数度量下训练的图表比基于方差或最优线性压缩的方法更能保持模型的预测。

英文摘要

How many dimensions does a language model's computation actually use? The question is ill-posed until one names a functional. Task-weighted charts make it well-posed: low-dimensional coordinate systems fit against a chosen functional of the representation, under the functional's own metric, turning distillation into plain least squares. Across six models from three families, spanning 70m to 7B parameters, next-token prediction needs 70--90% of the residual stream's width to stay within 5% of intact perplexity, a width consumed by the rare tail of language, and the variance profile predicts none of it: two directions carry 90% of GPT-2's activation variance and almost none of its function. Dimension is per-functional: the model's own uncertainty reads from six coordinates where the full predictive distribution needs hundreds; and it grows with depth. The dissociation is exploitable: when only a few dimensions can be kept, charts trained under the functional's metric preserve the model's predictions better than variance-based or optimal linear compression.

发表机构

  • Hother

机构由 AI 辅助整理,请以论文原文为准。

↑