arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10114cs.CL

长上下文混合模型机制 第1.1部分:从混合注意力到混合位置

Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position

Xiaoran Liu, Ziwei He, Xipeng Qiu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出长上下文混合模型机制,分析全注意力与SWA/LA混合的跷跷板效应,揭示SWA陷阱,并提出滑动窗口线性注意力实现16倍免训练外推且保持高准确率。

中文摘要 AI 辅助

大语言模型(LLM)的架构设计正从传统的纯全注意力模型转向混合模型,这类模型结合不同的注意力模块,以提高长上下文效率,并在长度外推和上下文扩展方面提升性能。为了解释混合模型为何有效以及如何更好地设计它们,我们提出了“长上下文混合模型机制”。作为本系列的第1.1部分,我们首先研究全注意力与滑动窗口注意力(SWA)或线性注意力(LA)的门控变体(以GLA和GDN为代表)的混合。我们首先观察到上下文扩展中的“跷跷板效应”:LA混合模型在长上下文持续预训练中获益更多,而SWA混合模型在长度外推下表现更好。我们将这种行为归因于这些注意力机制所引入的位置归纳偏置的差异。我们发现,SWA混合模型存在“短上下文学习陷阱”、“短窗口疲劳”和“长窗口懒惰”等问题,需要扩展窗口以增强其在持续长上下文预训练中的性能。对于LA混合模型,我们总结了“混合位置外推的马太效应”,并提出了滑动窗口线性注意力,实现了16倍的免训练长度外推,同时在64k上下文长度下保持NIAH-SK1上100%的准确率。

英文摘要

The architectural design of Large Language Models (LLMs) is shifting from traditional full-attention-only models to hybrid models, which combine different attention modules to improve long-context efficiency and performance in length extrapolation and context extension. To explain why hybrid models work and how to design them better, we propose Mechanics of Long-Context Hybrid Models. As Part 1.1 of this series, we begin with hybrids of full attention and either sliding-window attention (SWA) or gated variants of linear attention (LA), represented by GLA and GDN. We first observe a Seesaw Effect in Context Extension: LA hybrids benefit more from long-context continual pretraining, whereas SWA hybrids perform better under length extrapolation. We attribute this behavior to differences in the positional inductive biases induced by these attention mechanisms. We find that SWA hybrids suffer from a Short-Context Learning Trap, Short-Window Weariness, and Long-Window Laziness, and require extended windows to enhance performance in continual long-context pretraining. For LA hybrids, we summarize the Matthew Effect of Hybrid Position Extrapolation and propose Sliding-Window Linear Attention, achieving 16$\times$ training-free length extrapolation while maintaining 100\% accuracy on NIAH-SK1 in 64k context length.

发表机构

  • Shanghai Innovation Institute(上海创新研究院)
  • OpenMOSS Team(OpenMOSS团队)
  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑