arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LatentIndex:用于稀疏注意力的跨层共享与逐层选择

LatentIndex: Cross-Layer Sharing with Layer-Specific Selection for Sparse Attention

Zhaohui Wang, Zhixin Pan, Fanxu Meng, Muhan Zhang

arXiv 2610.04635首次发表:更新:

发表机构

Institute for Artificial Intelligence, Peking University; Dots Studio, Xiaohongshu Inc.; Beijing Institute of Technology(北京大学人工智能研究院; 小红书公司Dots工作室; 北京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LatentIndex通过跨层共享潜在缓存和逐层评分选择,减少稀疏注意力索引开销,在保持性能的同时显著降低缓存存储并提升解码速度。

AI 中文摘要

稀疏注意力减少了核心注意力计算,但其索引器仍会产生重复的选择工作和每层键缓存存储。跨层重用所选索引可减少此开销,但会将多个层限制为相同的令牌集。我们引入了LatentIndex,它将多头潜在注意力的潜在共享原理扩展到索引器层。每个层组从其第一层的隐藏状态构建共享潜在缓存,而逐层评分则实现独立的令牌选择。将键解码器吸收到查询中,可以直接对共享缓存进行评分,而无需重建历史逐层键。我们开发了免训练校准,并研究了该原理的训练感知实例化。为了平衡质量和计算,一种分层选择(HS)变体允许跟随者独立地细化由锚点提出的共享候选集。通过四层共享,LatentIndex在DeepSeek-V3.2上将逻辑索引器缓存存储减少了61.1%。在DeepSeek-V3.2和GLM-5上,免训练的LatentIndex在头部注意力质量召回率上比IndexCache提高了最多3.28个百分点,同时将RULER和LongBench性能保持在接近原生DSA的水平。HS在8K-128K上下文中进一步实现了相对于DSA的2.30-2.72倍解码索引器加速,同时保留了LatentIndex的大部分召回率。LatentIndex为跨层索引提供了新视角:共享连续表示而非离散选择,可在保持逐层令牌选择的同时实现高效重用。

英文摘要

Sparse attention reduces core-attention computation, but its indexers still incur repeated selection work and per-layer key-cache storage. Reusing selected indices across layers reduces this overhead but constrains multiple layers to the same token set. We introduce LatentIndex, which extends the latent-sharing principle of Multi-head Latent Attention across indexer layers. Each layer group constructs a shared latent cache from its first layer's hidden states, while layer-specific scoring enables independent token selection. Absorbing key decoders into queries enables direct scoring of the shared cache without reconstructing historical per-layer keys. We develop training-free calibration and investigate a training-aware instantiation of this principle. To balance quality and computation, a hierarchical selection (HS) variant lets followers independently refine a shared candidate set proposed by the anchor. With four-layer sharing, LatentIndex reduces logical indexer-cache storage by 61.1% on DeepSeek-V3.2. Across DeepSeek-V3.2 and GLM-5, training-free LatentIndex improves head-wise attention-mass recall over IndexCache by up to 3.28 percentage points while maintaining RULER and LongBench performance close to native DSA. HS further achieves 2.30-2.72 times decode indexer speedups over DSA across 8K-128K contexts, retaining most of LatentIndex's recall. LatentIndex offers a new perspective on cross-layer indexing: sharing continuous representations rather than discrete selections enables efficient reuse while preserving layer-specific token selection.

Commentspreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑