arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31261cs.CLcs.AI

MoSAR:用于学习自适应且可近似注意力几何的语义注意力机制混合

MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries

Michele Paolicelli, Alessandro Petruzzelli, Alessandro Franceso Maria Martina, Cataldo Musto, Giovanni Semeraro

首次发表
浏览论文内容

中文总结 AI 辅助

MoSAR通过输入条件路由器学习自适应、可控衰减的注意力几何,在5亿参数预训练中提升困惑度并支持top-1离散化,实现高效长上下文建模。

中文摘要 AI 辅助

密集自注意力的二次复杂度仍然是长上下文语言建模的核心瓶颈。许多高效的替代方案通过预先决定注意力应在何处稀疏或局部化来解决这一成本问题。我们认为,注意力近似应被视为一个几何问题,相关的交互几何应从数据中学习:自然语言依赖是输入相关的,难以预先规定,因此模型应学习在何处位置相关性可以衰减,在何处必须保留更广泛的交互。我们引入了语义注意力机制混合(MoSAR),它学习了查询-键交互上的这种自适应、可控衰减几何。输入条件的查询和键路由器在位置编码后应用,选择短程、中程和全局机制的混合,产生连续的、距离依赖的注意力场,而非固定的稀疏模式。这种几何在训练期间学习,随后可通过top-1路由进行离散化。在匹配的5亿参数模型的受控预训练实验中,MoSAR学习了显著更低覆盖范围的注意力几何,而不降低语言建模质量,在训练上下文长度上优于密集RoPE的困惑度。在长度外推下,MoSAR在所有评估变体中取得了最佳困惑度,包括ALiBi等强基线。此外,学习到的几何在确定性top-1离散化下保持稳定,表明它不仅自适应,而且在推理时也适用于低成本近似。

英文摘要

The quadratic complexity of dense self-attention remains a central bottleneck for long-context language modeling. Many efficient alternatives address this cost by deciding in advance where attention should be sparse or local. We argue that attention approximation should instead be approached as a geometric problem, with the relevant interaction geometry learned from data: natural-language dependencies are input-dependent and difficult to prescribe in advance, so the model should learn where positional relevance can decay and where broader interactions must be preserved. We introduce Mixture of Semantic Attention Regimes (MoSAR), which learns such an adaptive, controlled-decay geometry over query--key interactions. Input-conditioned query and key routers, applied after positional encoding, select mixtures over short, medium, and global regimes, inducing a continuous distance-dependent attention field rather than a fixed sparsity pattern. This geometry is learned during training and can subsequently be discretized through top-1 routing. In controlled pre-training experiments with matched 500M-parameter models, MoSAR learns a substantially lower-reach attention geometry without degrading language-modeling quality, improving perplexity over dense RoPE at the training context length. Under length extrapolation, MoSAR achieves the best perplexity among all evaluated variants, including strong baselines such as ALiBi. Moreover, the learned geometry remains stable under deterministic top-1 discretization, suggesting that it is not only adaptive, but also amenable to low-cost approximation at inference time.

发表机构

  • Università degli Studi di Bari Aldo Moro(巴里大学)

机构由 AI 辅助整理,请以论文原文为准。

↑