arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19203cs.CLcs.AIcs.LG

非对称注意力头:面向Transformer注意力的结构化头级上下文分配

Asymmetric Attention Heads: Structured Head-Wise Context Allocation for Transformer Attention

Zimu Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出非对称注意力头(AAH)框架,通过分层分组头并分配不同因果局部窗口,在4096-token实验中实现比纯全注意力更低的验证损失,为Transformer注意力提供结构化上下文分配机制。

中文摘要 AI 辅助

标准多头注意力(MHA)为每个注意力头提供相同的完整因果上下文范围,尽管不同头可承担不同的上下文角色:部分头主要依赖附近词汇或句法上下文,另一些头则依赖实体交互、话语关联或状态变化等长程关系。本文提出非对称注意力头(AAH),这是一种将上下文长度作为显式的头级或组级分配变量的头级上下文分配框架。AAH利用特征派生统计量对头进行分组,按层级组织这些组,并分配因果局部窗口,同时保留标准的平面MHA输出接口。在4096-token的seed-0实验中,多种AAH式局部分配变体实现了比纯全注意力更低的验证损失。短预算消融实验显示,稳定的局部分配和头窗口分配结构至关重要,而固定/局部控制可与自适应层级表现相当。我们将AAH解释为一种用于质量评估与分析的结构化头级上下文分配机制,注意力覆盖率(ACR)作为选定窗口路由诊断指标被报告。

英文摘要

Standard multi-head attention (MHA) gives every head the same full causal context span, although heads can serve different contextual roles. Some heads may rely mainly on nearby lexical or syntactic context, while others may depend on longer-range relations such as entity interactions, discourse links, or state changes. We present Asymmetric Attention Heads (AAH), a head-wise context- allocation framework that treats context length as an explicit per-head or per-group allocation variable. AAH groups heads using feature-derived statistics, organizes these groups hierarchically, and assigns causal local windows while preserving the standard flat MHA output interface. In 4096- token seed-0 experiments, several AAH-style local-allocation variants achieve lower validation loss than pure full attention. Short-budget ablations show that stable local allocation and head-window assignment structure matter, while fixed/local controls can be competitive with adaptive hierarchy. We interpret AAH as a structured head-wise context-allocation mechanism for quality and analysis, with Attention Coverage Ratio (ACR) reported as a selected-window routing diagnostic

补充信息

↑