arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

混合注意力大语言模型中的多语言性

Multilinguality in Hybrid Attention LLMs

Lucas Bandarkar, Junlin Hu, Chenyuan Yang, Mohsen Fayyaz, Nanyun Peng

arXiv 2609.35378首次发表:更新:

发表机构

University of California, Los Angeles; Fudan University(加州大学洛杉矶分校; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次探讨混合注意力架构对大语言模型多语言性的影响,发现跨语言对齐在首个全注意力层出现峰值,并通过蒸馏实验证明以全注意力层起始的层排序优于传统设计,学习速度提升达2.5倍。

AI 中文摘要

为应对智能体与推理用例中对长序列日益增长的需求,许多最先进的大语言模型结合多种注意力变体,以缓解传统softmax注意力的二次复杂度。这些混合注意力大语言模型旨在平衡全注意力与基于循环的替代方案各自的优势与局限。本工作首次研究了混合注意力如何影响大语言模型的多语言性。除了对分词不佳语言中长序列的影响外,我们的研究还受到以下可能性的驱动:循环状态的归纳偏置可能改变语言处理方式。我们的可解释性分析证实了这一点,表明混合模型中跨语言表征的发展模式与循环层和全注意力层的顺序相关。在多种模型中,我们尤其观察到在第一个全注意力层附近跨语言对齐出现显著峰值。这些发现促使我们质疑注意力层的传统排序。在多语言数据的蒸馏实验中,所有替代层排序在整个训练过程中均优于标准排序,学习速度最高提升2.5倍。这些显著且可复现的结果促使我们提出理论:多语言模型应始于全注意力层而非循环层。

英文摘要

In response to the growing demand for long sequences in agentic and reasoning use cases, many state-of-the-art LLMs combine multiple variants of attention to mitigate the quadratic complexity of traditional softmax attention. These hybrid attention LLMs aim to balance the strengths and limitations of full attention and alternatives based on recurrence. This work presents a first study of how hybrid attention impacts the multilinguality of LLMs. Beyond the impact on long sequences in poorly tokenized languages, our study is motivated by the possibility that the inductive biases of the recurrent state alter linguistic processing. Our interpretability analysis confirms this, showing that cross-lingual representations in hybrid models develop in patterns tied to the ordering of recurrent and full-attention layers. Across diverse models, we notably observe a pronounced spike in cross-lingual alignment around the first full-attention layer. These findings lead us to question the conventional ordering of attention layers. In distillation experiments on multilingual data, all alternative layer orderings outperform the standard throughout training, learning up to 2.5X faster. These stark, replicable results prompt our theory that multilingual models would benefit from starting with a full-attention layer rather than recurrent layers.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑