arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更快的查询-键学习强化自注意力模型中的注意力机制

Faster Query-Key Learning Sharpens Attention in Self-Attention Models

Rahul Vashisht, Harish G. Ramaswamy

arXiv 2608.06776首次发表:更新:

AI 中文总结

该研究通过分析自注意力模型的查询-键与输出-值回路的参数化,发现更快的查询-键学习可强化注意力、提升可解释性,且预测性能相当。

AI 中文摘要

标准自注意力层包含两个相互作用的回路:控制注意力分配的查询-键(query-key)回路,以及将被关注的表征映射到预测结果的输出-值(output-value)回路。查询-键与输出-值回路的折叠式(collapsed)和因式分解式(factorized)参数化会产生性质不同的注意力模式,其中部分参数化方式在训练损失相近的情况下,能对任务相关的token产生更锐利的注意力。我们分析了这些回路的参数化如何塑造用于下一个token预测训练的单层自注意力模型中的参数轨迹,通过梯度流分析表明,因式分解会诱导两个回路学习率的隐式重缩放。我们推导了闭式动力学,显示输出-值和查询-键参数沿一条直线移动,其相对速度由各自的学习率决定。因此,相对于输出-值学习更快的查询-键学习会产生更锐利的注意力,因为模型会通过增加对相关token的注意力权重来补偿输出-值学习的较慢速度。实验表明,两个回路的相对学习率差异决定了注意力的集中程度,这在保持可比预测性能的同时,提升了注意力可解释性代理指标。

英文摘要

A standard self-attention layer consists of two interacting circuits: the query-key circuit that governs attention allocation, and the output-value circuit that maps attended representations to predictions. Collapsed and factorized parameterizations of the query-key and output-value circuits lead to qualitatively different attention patterns. In particular, some parameterizations give sharper attention to task-relevant tokens, at a similar training loss. We analyze how the parameterizations of these circuits shape the parameter trajectories in single-layer self-attention models trained for next-token prediction. Through gradient-flow analysis, we show that factorization induces implicit rescaling of the two circuits' learning rates. We derive closed-form dynamics showing that output-value and query-key parameters move along a line, with relative speeds determined by their learning rates. Faster query-key learning relative to output-value learning thus produces sharper attention, as the model compensates for slower output-value learning by increasing attention mass on relevant tokens. Experiments show that differences in the relative learning rates of the two circuits govern attention concentration. This improves attention interpretability proxies while maintaining comparable predictive performance.

CommentsAccepted to the 43rd International Conference on Machine Learning (ICML 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑