arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00418cs.LG

CommunityKV:通过图分区实现高效长上下文解码

CommunityKV: Efficient Long-Context Decoding via Graph Partitioning

Joe McKenna, Anastasios Alexandridis, Nathan Susanj, Jing Liu

首次发表
浏览论文内容

中文总结 AI 辅助

针对长上下文解码中自注意力二次成本和KV缓存传输瓶颈,提出基于图划分的CommunityKV框架,将稀疏注意力视为社区检测,实现高效流式检索,吞吐量提升达1.71倍。

中文摘要 AI 辅助

将Transformer扩展到长上下文受到自注意力二次成本以及键值缓存内存传输线性增长的限制。稀疏注意力通过仅检索相关令牌来缓解这一问题,但当前方法要么需要大规模训练,要么在免训练场景下依赖语义粗糙的启发式方法或难以在解码过程中高效更新的昂贵聚类。我们提出了CommunityKV,一个将稀疏注意力表述为社区检测问题的框架。CommunityKV从标准预填充期间已计算的$QK^T$分数构建令牌图,并将图划分为社区,从而能够检索语义连贯的令牌组。一个局部更新规则以常数时间将新生成的令牌分配到社区,使得在整个流式解码过程中无需全局重新分区即可进行稀疏检索。我们在Qwen3和Llama-3.1模型上,跨三个长上下文基准评估了CommunityKV。每个查询头使用一个图时,CommunityKV相比密集注意力可实现高达$1.25\times$的端到端生成吞吐量,而查询组图聚合可实现高达$1.71\times$的吞吐量,同时保持相当的准确性。

英文摘要

Scaling Transformers to long contexts is constrained by the quadratic cost of self-attention and the linear growth of key-value cache memory transfer. Sparse attention mitigates this by retrieving only relevant tokens, but current approaches either require large-scale training or, within the training-free regime, rely on semantically coarse heuristics or expensive clustering that is difficult to update efficiently during decoding. We introduce CommunityKV, a framework that formulates sparse attention as a community detection problem. CommunityKV constructs a token graph from the $QK^T$ scores already computed during standard prefill, and partitions the graph into communities to enable retrieval of semantically coherent token groups. A local update rule assigns newly generated tokens to communities in constant time, enabling sparse retrieval throughout streaming decoding without global re-partitioning. We evaluate CommunityKV on Qwen3 and Llama-3.1 models across three long-context benchmarks. With one graph per query head, CommunityKV delivers up to $1.25\times$ the end-to-end generation throughput of dense attention, while query-group graph aggregation yields up to $1.71\times$ with comparable accuracy.

发表机构

  • Amazon AGI(亚马逊AGI)

机构由 AI 辅助整理,请以论文原文为准。

↑