arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多头注意力残差

Multi-Head Attention Residuals

Cheng Luo, Zefan Cai, Junjie Hu

arXiv 2607.27230首次发表:更新:

发表机构

University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出多头注意力残差(MHAR),通过多子空间头优化注意力残差,在多参数规模下降低Transformer验证损失,提升训练吞吐量,中途训练8B模型在GSM8K和GPQA上获显著增益。

AI 中文摘要

Transformer通过单一的加法残差流在深度维度传播信息:每个子层仅读取最新状态。注意力残差通过让每个子层通过学习到的softmax进行注意力读取,放松了这一限制。然而,这种读取使用了跨整个宽度共享的单一查询,因此每个特征子空间必须通过单一分布读取深度历史。这种强制妥协的成本随子空间对应读取哪些层的分歧程度增加而上升,而分歧程度随模型宽度增大而增加。我们引入多头注意力残差(Multi-Head Attention Residuals, MHAR):将路由查询重塑为H个各子空间专用的头,每个头对深度历史拥有各自的softmax。这种读取变为块对角形式,重塑操作不增加参数,计算开销可忽略,且当H=1时可完全恢复注意力残差。在基于去重Nemotron的退火语料库(经质量过滤且以STEM和代码内容为主)上从头训练后,MHAR在1亿、3.5亿和10亿参数规模下,相比标准Transformer的验证损失分别降低了-0.061、-0.149和-0.140。它在所有设置下均取得四种方法中的最佳结果,且增益从1亿参数到更大规模逐步提升。头数是一个实际设计维度而非自由旋钮:验证损失关于H呈U形,各规模下在H=4或H=8处达到平稳最优。我们在大规模模型中采用H=8;超过该点过度拆分(H=16)会持续损失部分增益。对训练后查询的直接探测证实,学习到的子空间分歧是背后的驱动因素。融合Triton路由内核将注意力残差训练吞吐量从基线的0.2-0.5倍提升至0.55-0.88倍,同时保持接近基线的峰值内存。使用delta注意力残差的保身份转换支持80亿参数模型的中途训练,在GSM8K上提升3.2,在GPQA上提升3.1。

英文摘要

Transformers propagate information across depth through a single additive residual stream: every sublayer reads only the most recent state. Attention residuals relax this by letting each sublayer attend, through a learned softmax. However, that read uses a single query shared across the entire width, so every feature subspace must read the depth history through one distribution. The cost of this forced compromise grows with how much the subspaces disagree about which layers to read, and disagreement grows with model width. We introduce Multi-Head Attention Residuals (MHAR): the routing query is reshaped into H per-subspace heads, each with its own softmax over the depth history. The read becomes block-diagonal, the reshape adds zero parameters and negligible compute, and H = 1 recovers attention residuals exactly. Trained from scratch on a deduplicated Nemotron-based anneal corpus that is quality-filtered and STEM- and code-heavy, MHAR improves validation loss over a standard Transformer at 100M, 350M, and 1B (-0.061, -0.149, and -0.140). It achieves the best result among four methods in every setting, with the gain increasing from 100M to the larger scales. The head count is a real design axis rather than a free knob: validation loss is U-shaped with respect to H, with a flat optimum at H = 4 or H = 8 across scales. We adopt H = 8 for large-scale models; over-splitting beyond this point (H = 16) consistently gives back part of the gain. A direct probe of the trained queries confirms that learned subspace disagreement is the underlying driver. Fused Triton routing kernels increase attention-residual training throughput from 0.2-0.5x to 0.55-0.88x of the baseline while maintaining near-baseline peak memory. An identity-preserving conversion using delta attention residuals supports 8B mid-training, yielding improvements of +3.2 on GSM8K and +3.1 on GPQA.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑