发表机构
LLM-Core Xiaomi(小米LLM-Core)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HySparse2通过两级KV共享(外层KV桥接与内层令牌级KV重用)实现高效预填充和紧凑缓存,在80B-A3B MoE模型上提升长上下文检索与多轮智能体任务性能。
AI 中文摘要
长时程和多轮智能体通常生成简短的动作,并处理来自工具和环境的长时间观察。这种不断增长的上下文要求高效的预填充、紧凑的KV缓存存储以及准确的长上下文检索。为了满足这些需求,我们引入了HySparse2,一种具有两级KV共享的混合稀疏注意力架构。在外层,KV桥接采用YOCO风格的自解码器和交叉解码器结构,但仅桥接全注意力层。自解码器使用混合滑动窗口注意力(SWA),而交叉解码器使用混合稀疏注意力。交叉解码器中全注意力层的KV缓存由自解码器中全注意力层的隐藏状态生成。在内层,HySparse2保留了HySparse的核心KV重用设计,并进行了两项改进。首先,它将块级稀疏性替换为令牌级稀疏性,以实现更精细的长上下文检索。其次,它从稀疏层中移除单独的SWA分支,而是将最近令牌的滑动窗口强制纳入稀疏选择中。这种两级KV共享允许所有交叉解码器KV缓存从自解码器隐藏状态构建。因此,预填充可以在自解码器之后退出,跳过所有交叉解码器层。在80B-A3B MoE模型上,HySparse2在长上下文检索和多轮智能体任务上优于HySparse和混合SWA,同时大幅减少预填充计算和KV缓存存储。
英文摘要
Long-horizon and multi-turn agents typically generate short actions and process long observations from tools and environments. This growing context demands efficient prefill, compact KV-cache storage, and accurate long-context retrieval. To meet these demands, we introduce HySparse2, a hybrid sparse attention architecture with two-level KV sharing. At the outer level, KV Bridging adopts a YOCO-style self-decoder and cross-decoder structure, but bridges only full-attention layers. The self-decoder uses hybrid sliding-window attention (SWA), while the cross-decoder uses hybrid sparse attention. The KV caches for full-attention layers in the cross-decoder are generated from the hidden states of full-attention layers in the self-decoder. At the inner level, HySparse2 retains HySparse's core KV Reuse design with two refinements. First, it replaces block-level sparsity with token-level sparsity for finer long-context retrieval. Second, it removes the separate SWA branch from sparse layers and instead forces a sliding window of recent tokens into the sparse selection. This two-level KV sharing allows all cross-decoder KV caches to be constructed from self-decoder hidden states. Prefill can therefore exit after the self-decoder, skipping all cross-decoder layers. On an 80B-A3B MoE model, HySparse2 outperforms HySparse and Hybrid SWA on long-context retrieval and multi-turn agentic tasks, while substantially reducing prefill computation and KV-cache storage.