arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

激活稀疏化与KV缓存稀疏化在LLM解码中的交叉点

Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding

Jungseob Lee, Seungyoon Lee, Seongtae Hong, Sugyeong Eo, Heuiseok Lim

arXiv 2609.33889首次发表:更新:

发表机构

Korea University; Yonsei University Mirae Campus(高丽大学; 延世大学未来校区)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推导出激活稀疏化与KV缓存稀疏化在LLM解码中的字节交叉点,并实验验证了组合策略在长上下文下的速度提升,最高达26%。

AI 中文摘要

在每一步解码中,使用大型语言模型解码一个序列会重新读取投影权重,其流量是固定的,以及键值(KV)缓存,其流量随上下文增长。激活稀疏化削减了第一项,KV缓存稀疏化削减了第二项,然而它们报告的速度提升难以比较,因为每项都依赖于上下文长度以及与之比较的密集注意力内核。我们推导出字节交叉点,即两项节省相等时的上下文长度,以及每个分支及其组合的理想速度提升界限,仅基于模型维度和保留比率。然后,我们在两个GPU上,在真实文本的密集预填充之后,对2K到128K令牌范围内的两个分支及其组合进行计时,密集和稀疏模式通过相同的split-K注意力内核读取缓存。投影分支在短上下文时领先,KV分支在长上下文时领先,其速度提升遵循其字节界限直至固定内核成本。在单独扫描中测量的这些成本,使得字节账户能够预测三个保留比率对、第二个模型和第二个GPU的测量交叉点,误差在4.1K令牌以内。使用掩蔽而非split-K注意力对密集基线进行计时,会使相同KV策略的表观速度提升膨胀约五倍。一种注意力评分的KV选择在高达127K令牌的范围内,与密集解码一样回答相同的密码和多键放置,而KV窗口则遗漏了大部分。在匹配的困惑度预算下,激活稀疏化与此选择组合,在两个GPU上的解码速度比最佳单分支快14%至26%。代码可在https URL获取。

英文摘要

At each step, decoding one sequence with a large language model rereads the projection weights, whose traffic is fixed, and the key-value (KV) cache, whose traffic grows with context. Activation sparsity trims the first term and KV-cache sparsity the second, yet their reported speedups are hard to compare because each depends on context length and on the dense attention kernel it is measured against. We derive a byte crossover, the context length at which the two savings are equal, together with ideal speedup bounds for each branch and for their composition, from model dimensions and keep ratios alone. We then time both branches and their composition from 2K to 128K tokens on two GPUs after a dense prefill of real text, with dense and sparse modes reading the cache through the same split-K attention kernel. The projection branch leads at short context and the KV branch at long context, with speedups that follow their byte bounds up to fixed kernel costs. Adding these costs, measured in separate sweeps, lets the byte account predict the measured crossings of three keep-ratio pairs, a second model, and a second GPU to within 4.1K tokens. Timing the dense baseline with masked instead of split-K attention inflates the apparent speedup of the same KV policy about fivefold. An attention-scored KV selection answers the same passkey and multi-key placements as dense decoding up to 127K tokens, whereas a KV window misses most of them. Under matched perplexity budgets, activation sparsity composed with this selection decodes 14 to 26% faster than the best single branch on both GPUs. Code is available at https://github.com/js-lee-AI/ByteCross.

Comments22 pages, 6 figures, 17 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑