arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HySparse2:具有两级KV共享的混合稀疏注意力

HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing

Jianyu Wei, Yizhao Gao, Qihao Zhang, Shimao Chen, Zhengju Tang, Yu Cheng, Shengjie Zhou, Zihan Jiang, Yifan Song, Hailin Zhang, Liang Zhao, Bo Yang, Gang Wang, Shijie Cao, Fuli Luo

arXiv 2609.26368首次发表:更新:

发表机构

LLM-Core Xiaomi(小米LLM-Core)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

HySparse2通过两级KV共享(外层KV桥接与内层令牌级KV重用)实现高效预填充和紧凑缓存,在80B-A3B MoE模型上提升长上下文检索与多轮智能体任务性能。

AI 中文摘要

长时程和多轮智能体通常生成简短的动作,并处理来自工具和环境的长时间观察。这种不断增长的上下文要求高效的预填充、紧凑的KV缓存存储以及准确的长上下文检索。为了满足这些需求,我们引入了HySparse2,一种具有两级KV共享的混合稀疏注意力架构。在外层,KV桥接采用YOCO风格的自解码器和交叉解码器结构,但仅桥接全注意力层。自解码器使用混合滑动窗口注意力(SWA),而交叉解码器使用混合稀疏注意力。交叉解码器中全注意力层的KV缓存由自解码器中全注意力层的隐藏状态生成。在内层,HySparse2保留了HySparse的核心KV重用设计,并进行了两项改进。首先,它将块级稀疏性替换为令牌级稀疏性,以实现更精细的长上下文检索。其次,它从稀疏层中移除单独的SWA分支,而是将最近令牌的滑动窗口强制纳入稀疏选择中。这种两级KV共享允许所有交叉解码器KV缓存从自解码器隐藏状态构建。因此,预填充可以在自解码器之后退出,跳过所有交叉解码器层。在80B-A3B MoE模型上,HySparse2在长上下文检索和多轮智能体任务上优于HySparse和混合SWA,同时大幅减少预填充计算和KV缓存存储。

英文摘要

Long-horizon and multi-turn agents typically generate short actions and process long observations from tools and environments. This growing context demands efficient prefill, compact KV-cache storage, and accurate long-context retrieval. To meet these demands, we introduce HySparse2, a hybrid sparse attention architecture with two-level KV sharing. At the outer level, KV Bridging adopts a YOCO-style self-decoder and cross-decoder structure, but bridges only full-attention layers. The self-decoder uses hybrid sliding-window attention (SWA), while the cross-decoder uses hybrid sparse attention. The KV caches for full-attention layers in the cross-decoder are generated from the hidden states of full-attention layers in the self-decoder. At the inner level, HySparse2 retains HySparse's core KV Reuse design with two refinements. First, it replaces block-level sparsity with token-level sparsity for finer long-context retrieval. Second, it removes the separate SWA branch from sparse layers and instead forces a sliding window of recent tokens into the sparse selection. This two-level KV sharing allows all cross-decoder KV caches to be constructed from self-decoder hidden states. Prefill can therefore exit after the self-decoder, skipping all cross-decoder layers. On an 80B-A3B MoE model, HySparse2 outperforms HySparse and Hybrid SWA on long-context retrieval and multi-turn agentic tasks, while substantially reducing prefill computation and KV-cache storage.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑