arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

软注意力中的弃权与噪声过滤:两个缺失的原语

Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

Richard Zhe Wang

arXiv 2609.22005首次发表:更新:

发表机构

St. John Fisher University(圣约翰费舍尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出并验证了软注意力缺失的两种原语——弃权(不执行)与噪声过滤,发现其收益随模型规模呈相反趋势,且两者结合在10M至350M参数规模上均达到最优性能。

AI 中文摘要

据报道,对注意力值通路进行门控可以改善语言模型的预训练,而先前的研究对其原因存在分歧。我们提出并提供了实验证据,表明此类门控提供了软注意力所缺乏的两种不同功能:弃权(不执行)和噪声过滤。第一种是弃权(不执行),它允许注意力头不输出任何内容,从而绕过了注意力权重之和必须为一的要求。第二种是噪声过滤,它允许注意力头的值通路抑制残差流中叠加特征的干扰。在我们从10M到350M参数的匹配模型实验中,我们通过softmax中学习到的每头汇点逻辑提供弃权(不执行),并通过每个值上的门控提供噪声过滤。我们报告了三个实证发现。首先,弃权(不执行)的益处(以相对于匹配基线的验证损失减少量衡量)随着模型规模的增大而下降,而噪声过滤的益处则随规模增大而增加。特别是,在10M规模时,弃权(不执行)几乎占据了门控带来的全部收益,而在350M规模时,过滤则占据了大部分收益。其次,在每个规模上,最佳模型是同时内置这两种原语的模型。第三,向注意力头读取的值中注入受控干扰,证实了门控能够消除此类干扰,并揭示了我们所研究的两种门控形式各自具有一个特征性的盲点。提供这两种原语仅增加可忽略不计的参数,并且与键值缓存保持兼容。

英文摘要

Softmax attention has two structural gaps. A head cannot abstain, because its weights sum to one, so it outputs something even when nothing is relevant. Nor can it filter what it reads, because its output is a weighted average of value vectors, passing interference as faithfully as signal. We call these missing primitives abstention and noise filtering. Recent studies report that gating the value pathway improves pretraining but attribute the gain to different causes. We show that a value gate partly supplies both primitives, which unifies the reported causes as views of one gain. We give each primitive its own mechanism in matched models of 10M to 350M parameters and measure what each contributes. The gain from gating is almost entirely abstention at 10M, whereas by 350M filtering contributes as much as abstention, so what a study observes depends on its scale. The two benefits are largely additive, with a small overlap. A gate determined by each value alone leaves the attention sink in place, whereas a query-controlled mechanism removes it. Injecting interference into the value reads shows that abstention and filtering protect against it in distinguishable ways. The same patterns appear in pretrained models up to 20B parameters.

Comments20 pages (8 pages main text plus appendices), 5 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑