发表机构
AIGCode; South China University of Technology; Pazhou Laboratory(AIGCode; 华南理工大学; 琶洲实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM预训练效率问题,提出含局部融合注意力与知识记忆模块的LoKiFormer,使预训练收敛速度提升1.33倍,性能优于现有架构。
AI 中文摘要
大语言模型(LLMs)已在各类应用中取得显著突破,但其架构在预训练阶段仍存在效率不足的问题,主要源于两大局限:(i)自注意力缺乏明确的局部性归纳偏置,导致对序列内部局部信息的冗余建模;(ii)混合专家(MoE)隐式地将知识存储与计算路径耦合,阻碍了对序列外部全局知识的灵活访问。为克服这些局限,本文提出LoKiFormer——一种新型LLM架构,在标准解码器基础上新增两个专用模块:1)局部融合注意力(LFA),将卷积融合引入注意力机制,明确捕捉局部模式,使注意力能作用于更具信息性的表示;2)知识记忆模块(KMM),引入参数化键值记忆,在可寻址槽中显式存储全局知识,实现存储与计算的解耦,支持直接知识检索。这些模块协同使LoKiFormer能在两个层面实现更高效、更有效的信息整合。实验结果显示,LoKiFormer的预训练收敛速度比基线模型快1.33倍,凸显其优于现有LLM架构的优势。
英文摘要
Large language models (LLMs) have achieved remarkable breakthroughs across various applications. However, their architectures remain inefficient in pretraining due to two main limitations: (i) self-attention lacks an explicit inductive bias for locality, leading to redundant modeling of sequence-internal local information; (ii) mixture-of-experts (MoE) implicitly couples knowledge storage with computational pathways, hindering flexible access to sequence-external global knowledge. To overcome these limitations, we propose LoKiFormer, a novel LLM architecture that augments the standard decoder with two dedicated modules: 1) Local Fusion Attention (LFA), which incorporates a convolutional fusion to attention, explicitly capturing local patterns and allowing the attention to operate on more informative representations; 2) Knowledge Memory Module (KMM), which introduces a parametric key-value memory that explicitly stores global knowledge in addressable slots, decoupling storage from computation and enabling direct knowledge retrieval. Together, these modules enable LoKiFormer to achieve more efficient and effective integration of information at both levels. Experimental results show that LoKiFormer converges 1.33x faster in pre-training than baseline models, underscoring its superiority over existing LLM architectures.
CommentsAccepted by ICML 2026