arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06849cs.CLcs.AI

头自主性:基于冻结查询-键几何的无数据稀疏注意力

Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry

  • Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
  • University of Chinese Academy of Sciences(中国科学院大学)
  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • Baidu Inc.(百度公司)

机构由 AI 辅助整理,请以论文原文为准。

Yehan Yang, Junyuan Shang, Yang Li, Guanqun Zhao, Shuohuan Wang, Dianhai Yu

AI总结:

本研究提出无数据稀疏注意力方法 AoH,通过查询-键投影谱几何识别检索头与流头,在 50% 稀疏度下平均保留全注意力 96.5% 性能,同时显著降低 LLM 的计算延迟与 KV-cache 内存。

AI中文摘要:

长上下文大语言模型(LLM)推理受限于二次注意力计算和不断增长的键值缓存(KV-cache)成本。现有稀疏注意力和键值压缩方法通常从运行时注意力分数、观测窗口、校准提示或学习门控中决定保留哪些 token 或头,这使得头诊断依赖输入且部署成本高昂。我们提出了头自主性(Autonomy-of-Heads,AoH),一种无数据方法,可从查询-键投影的谱几何中识别检索头和流头。AoH 定义了核注意力算子 $M_h = W_K^{h\ op}W_Q^h$,并将其有效秩作为头功能的权重空间度量:集中的谱表明存在少量主导的查询-键匹配方向,与检索头相关;而弥散的谱表明不存在主导的全局匹配方向,与流头相关。我们进一步推导了高效的 $d_{\ ext{head}}$ 维计算,避免构建完整的 $d_{\ ext{model}}\ imes d_{\ ext{model}}$ 矩阵。我们在多个模型上进行了广泛实验,结果表明,在 50% 稀疏度下,AoH 平均保留了全注意力 96.5% 的性能,同时将预填充延迟和解码延迟分别降低了多达 41.4% 和 66.0%,在 256K token 时将 KV-cache 内存降低了 50.0%。

英文摘要:

Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs. Existing sparse attention and KV-compression methods typically decide which tokens or heads to preserve from runtime attention scores, observation windows, calibration prompts, or learned gates, making head diagnosis input-dependent and costly to deploy. We propose Autonomy-of-Heads (AoH), a data-free method that identifies retrieval and streaming heads from the spectral geometry of query-key projections. AoH defines the kernel attention operator $M_h = W_K^{h\top}W_Q^h$ and uses its effective-rank as a weight-space measure of head function: concentrated spectra indicate a small number of dominant query-key matching directions and are associated with retrieval heads, whereas diffuse spectra indicate the absence of a dominant global matching direction and are associated with streaming heads. We further derive an efficient $d_\text{head}$-dimensional computation that avoids constructing the full $d_\text{model}\times d_\text{model}$ matrix. We conducted extensive experiments across models demonstrating that at 50\% sparsity, AoH retains 96.5\% of Full Attention performance on average while reducing prefill and decode latency by up to 41.4\% and 66.0\%, respectively, and KV-cache memory by 50.0\% at 256K tokens.

↑