发表机构
University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过直接分析Transformer的参数和激活,揭示预测依赖的组件数量远少于表面贡献,并展示如何利用这些发现进行高效编辑和安装。核心贡献在于无需训练即可读取和写入模型组件。
AI 中文摘要
一个Transformer的多少组件决定一个token?按每个单元和通道对logit贡献的绝对值计数,一次预测依赖于数千到数十万个组件。但贡献是有符号的,在十八个模型中,背离预测token的质量中位数是承载它的质量的七倍。除以净值后,计数降至几十个:在基线上,53个组件承载了预测的百分之九十,13个组件是预测无法承受失去的,而8个组件足以单独产生预测。在另外训练的十二个模型中,参数规模从1.24亿到70亿不等,充分集从两个组件到十六个不等,而预测所依赖的部分,一路追溯回去,占模型的百分之一到百分之三,这一比例不随规模增长。一层更新的四分之三是其所接收状态的固定线性映射。所有内容都从模型自身的参数和激活中读取,没有训练或拟合任何东西,并且它在两侧命名组件:它写入的内容,从其驱动的预测来看,接近每个模型的一半;它读取的内容,从其自身层框架中的权重来看,在其八个最强输入上比随机水平高出58.9%。按上游来源对剩余部分排序,可以得到嵌入层无法看到的语法类别。一个名字可以被操作。模型不持有的一种关联安装到一个备用单元中,键和值从权重中读取,代价是百分之零点二五的留出损失,是秩一更新代价的四十分之一。一个安装的注意力头及其上方两层的单元使得编辑仅在上下文中较早出现token的位置触发,而模型为自己训练的一个单元由上游两层驱动,86%的效果通过它传递。一个保序激活将单元的输入置于仪器的上限,代价是两部分安装。
英文摘要
How many of a transformer's components decide a token? Counted by the absolute value of each unit's and channel's contribution to the logit, one prediction rests on thousands to hundreds of thousands of them. But contributions are signed, and across eighteen models the mass pushing away from the predicted token is a median of seven times the mass carrying it. Divide by the net and the count is dozens: on the baseline, 53 components carry ninety percent of a prediction, 13 it cannot survive losing, and 8 suffice to produce it alone. Across twelve models trained elsewhere, 124M to 7B parameters, the sufficient set runs from two components to sixteen, and what a prediction draws on, followed all the way back, is one to three percent of the model, a share that does not grow with size. Three quarters of a layer's update is a fixed linear map of the state it received. Everything is read from the model's own parameters and activations, with nothing trained or fitted, and it names a component on both sides: what it writes, from the predictions it drives, reaching close to half of every model; what it reads, from its weights in the frame of its own layer, at 58.9 percent above chance over its eight strongest inputs. Sorting the remainder by upstream source yields grammatical categories the embedding cannot see. A name can be acted on. An association the model does not hold installs into one spare unit, key and value read from the weights, for a quarter of a percent of held-out loss, a fortieth of what a rank-one update costs. An installed attention head and a unit two layers above it make an edit fire only where a token occurred earlier in the context, and a unit the model trained for itself is driven from two layers upstream, 86 percent of the effect passing through it. An order-preserving activation puts a unit's inputs at the instrument's ceiling, at the price of a two-part install.