arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HLA:通过分块动态混合实现富有表现力的混合线性注意力

HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang

arXiv 2610.05842首次发表:更新:

发表机构

Monash University; Zhejiang University(莫纳什大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对线性注意力历史压缩导致稀疏信息访问困难的问题,提出查询依赖的分块级混合线性注意力HLA,通过仿射状态转移与内容相关路由门控动态组合历史,在Qwen3.5及从头训练中显著提升长上下文性能。

AI 中文摘要

线性注意力通过将历史信息压缩为循环状态,实现了高效的长上下文自回归解码,但这种压缩可能使对稀疏和遥远信息的选择性访问变得困难。现有的基于分块的扩展增加了记忆容量,但学习到的分块混合系数可能相对于输入内容保持固定,因此无法针对每个查询调整历史访问。我们提出了混合线性注意力(HLA),一种用于门控Delta网络(GDN)的查询依赖的分块级注意力机制。HLA将每个已完成的分块表示为精确的仿射状态转移,并从紧凑的、自注意力池化的代表中计算内容相关的路由门控。每个门控将相应的历史转移与恒等映射进行插值,同时控制该分块的加性记忆及其对早期状态的变换。有效支持正则化进一步鼓励稀疏推理的集中路由。我们在预训练适配和从头训练两种设置下评估了HLA。在从0.8B到9B的Qwen3.5模型中,HLA始终优于原生GDN和固定分块混合,在LongBench-V2上最高提升5.57个百分点,在RULER上提升3.97个百分点。在受控的从头训练1.3B设置中,使用4K上下文训练100B个token,HLA还将RULER性能从4K提升到32K,增益从4K时的0.83个百分点增加到32K时的4.22个百分点。这些结果表明,查询依赖的循环记忆组合改善了长上下文建模,并在超出训练上下文的情况下保持有效,同时使用紧凑的每分块仿射摘要。项目页面:此https URL

英文摘要

Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective access to sparse and distant information difficult. Existing chunk-based extensions increase memory capacity, yet learned chunk-mixing coefficients may remain fixed with respect to input content and therefore cannot adapt historical access to each query. We introduce \emph{Hybrid Linear Attention} (HLA), a query-dependent chunk-level attention mechanism for Gated DeltaNet (GDN). HLA represents each completed chunk as an exact affine state transition and computes content-dependent routing gates from compact, self-attentively pooled representatives. Each gate interpolates the corresponding historical transition with the identity map, controlling both the chunk's additive memory and its transformation of earlier states. Effective-support regularization further encourages concentrated routing for sparse inference. We evaluate HLA under both pretrained adaptation and from-scratch training. Across Qwen3.5 models from 0.8B to 9B, HLA consistently improves over native GDN and fixed chunk mixing, with gains of up to 5.57 percentage points on LongBench-V2 and 3.97 points on RULER. In a controlled from-scratch 1.3B setting trained for 100B tokens with a 4K context, HLA also improves RULER performance from 4K to 32K, with gains increasing from 0.83 points at 4K to 4.22 points at 32K. These results demonstrate that query-dependent composition of recurrent memory improves long-context modeling and remains effective beyond the training context while using compact per-chunk affine summaries. Project page: https://caesarhhh.github.io/hla/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑