发表机构
International Institute of Information Technology Hyderabad(海得拉巴国际信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MS-GLA通过多时间分辨率分解注意力头,解决GLA的表征瓶颈,在不增加参数下提升记忆容量,在语言建模和回忆任务上优于GLA。
AI 中文摘要
门控线性注意力(GLA)Transformer通过数据相关的门控机制推进了线性递归模型的发展,但面临一个核心限制:所有注意力头中的固定容量记忆矩阵以单一时间分辨率运行,每个token被单独处理,迫使它们同时编码局部句法模式和长距离语义结构,从而产生仅靠门控不足以解决的表征瓶颈。我们引入了多尺度门控线性注意力(MS-GLA),通过将注意力头分布到多个时间分辨率来解决这一问题。较粗的分辨率对更长的token跨度进行池化,自然地专门处理长距离依赖,而较细的头组则保持对局部句法结构的敏感性。一个可学习的、输入相关的融合层在每个时间步动态重组头组输出,在不增加每个头状态大小的情况下扩展有效记忆容量。这种多分辨率分解借鉴了多尺度状态空间模型(MS-SSM)的原理,并将其适应于门控线性注意力设置。我们在语言建模、回忆密集型任务和长上下文泛化上评估了MS-GLA。在所有设置中,MS-GLA在匹配参数数量下始终比GLA实现更高的准确率和更低的困惑度,在回忆密集型任务上最高提升18.9%,在语言建模基准上平均困惑度降低9.5%,验证了多时间分辨率分解作为门控线性注意力的一种有原则且有效的扩展。
英文摘要
Gated Linear Attention (GLA) Transformers advance linear recurrent models through data-dependent gating, but face a core limitation: the fixed-capacity memory matrices across all heads operate at a single temporal resolution, where each token is processed individually, forcing them to simultaneously encode local syntactic patterns and long-range semantic structure, creating a representational bottleneck that gating alone is insufficient to resolve. We introduce Multi-Scale Gated Linear Attention (MS-GLA), which addresses this by distributing attention heads across multiple temporal resolutions. Coarser resolutions pool longer token spans naturally specializing toward long-range dependencies, while finer head groups retain sensitivity to local syntactic structure. A learnable, input-dependent fusion layer dynamically recombines head group outputs at each timestep, expanding effective memory capacity without increasing per-head state size. This multi-resolution decomposition draws on principles from Multi-Scale State-Space Models (MS-SSM), adapting them to the gated linear attention setting. We evaluate MS-GLA on language modeling, recall-intensive tasks, and long-context generalization. Across all settings, MS-GLA consistently achieves higher accuracy and lower perplexity than GLA at matched parameter counts, with up to 18.9% improvement on recall-intensive tasks and 9.5% lower average perplexity on language modeling benchmarks, validating multi-temporal resolution decomposition as a principled and effective extension of Gated Linear Attention.