arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

穿越瓶颈:多头潜在注意力如何在语言模型中分离内容与位置

Through the Bottleneck: How Multi-head Latent Attention Separates Content from Position in Language Models

Dhruvil S, Fenil Sojitra, Ravirajsinh Chauhan

arXiv 2607.23054首次发表:更新:

发表机构

Indian Institute of Technology Madras; P P Savani University(印度理工学院马德拉斯分校; P P萨瓦尼大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多头潜在注意力(MLA)在语言模型中如何分离内容与位置。通过训练特定Transformer并运用多种分析方法,发现其瓶颈能保留内容丢弃位置信息,归纳头集中在单层,有“语义中心”层,且瓶颈容量供应过剩,重塑了模型相关结构。

AI 中文摘要

在DeepSeek-V2中引入的多头潜在注意力(MLA),通过共享的低秩瓶颈(cKV)压缩键值对,在推理过程中实现了81%的键值缓存减少。尽管它已被大规模生产模型采用,但此前没有工作研究过这个瓶颈保留或丢弃了哪些信息,以及它如何重塑内部Transformer电路。我们首次对MLA进行了全面的机制可解释性研究,训练了一个1.14亿参数的Transformer(在网页/代码/数学混合数据上预训练,在TinyStories上微调),并通过奇异值分解、注意力头分类、线性探测和干扰归因分析来分析其表示。我们的主要发现包括:cKV瓶颈学习到了一个纯粹的内容表示,保留了实体身份(保留率98%),同时丢弃了位置信息,验证了MLA通过旋转位置编码(RoPE)分离内容与位置的能力;归纳头集中在单层(第12层),与标准多头注意力(MHA)中的分散分布不同;单个“语义中心”层(第15层)同时表现出最高的奇异值分解有效秩和最强的干扰归因分数;瓶颈在全局上供应过剩,平均仅使用其46%的容量。这些发现表明,MLA不仅被动地压缩注意力,还重塑了模型组织内容、位置和电路结构的方式。我们将此视为一个初步的数据点,并在第5节中详细说明了范围限制。

英文摘要

Multi-head Latent Attention (MLA), introduced in DeepSeek-V2, compresses key-value pairs through a shared low-rank bottleneck (cKV), achieving 81% KV-cache reduction during inference. Despite its adoption in massive production models, no prior work has studied what information this bottleneck preserves or discards, nor how it reshapes internal transformer circuits. We present the first comprehensive mechanistic interpretability study of MLA, training a 114M-parameter transformer (pretrained on a web/code/math mixture, fine-tuned on TinyStories) and analyzing its representations through SVD, attention head taxonomy, linear probing, and a disruption-attribution analysis. Our key findings are: (1) the cKV bottleneck learns a pure content representation, preserving entity identity (98% retention) while discarding positional information, validating MLA's separation of content from position via RoPE; (2) induction heads co-locate at a single layer (Layer 12), unlike their distributed formation in standard MHA; (3) a single "semantic hub" layer (Layer 15) simultaneously exhibits the highest SVD effective rank and strongest disruption-attribution score; and (4) the bottleneck is globally over-provisioned, using only 46% of its capacity on average. These findings suggest MLA does not merely compress attention passively, but reshapes how the model organizes content, position, and circuit structure. We view this as an initial data point and detail scope limitations in Section 5.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑