arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于自适应注意力匹配的思维感知KV缓存压缩方法用于推理

Thought-Aware KV Cache Compaction for Reasoning via Adaptive Attention Matching

Yang Liu, Bin Chong, Chongyang Zhang, Hao Zheng, Jiayu Liang, Xu Kefu

arXiv 2608.12331首次发表:更新:

发表机构

Tsinghua University; Peking University; Soochow University(清华大学; 北京大学; 苏州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对推理语言模型 KV 缓存的内存瓶颈问题,提出思维感知注意力匹配(TAM)方法,通过三种机制实现高效压缩,在降低内存占用的同时保持了推理准确率。

AI 中文摘要

推理语言模型会生成长长的思维链(CoT)序列,其键值(KV)缓存呈线性增长,在解码过程中成为内存瓶颈。现有的压缩方法将推理轨迹视为扁平的 token 序列并应用统一压缩,忽略了 CoT 推理的分层结构,其中不同步骤的重要性差异极大。我们提出**思维感知注意力匹配(TAM)**,该方法通过三种机制利用这一结构:(i)将轨迹分解为推理块的思维分段;(ii)基于每个分段的重要性和大小分配压缩预算的自适应预算分配;(iii)保留高注意力推理锚点的关键 token 保护。我们证明,在凸误差模型下该分配规则是最优的,且顺序压缩下的累积误差保持有界。在 AIME 2024 和 MATH-500 数据集上使用 Qwen3-4B 模型的实验表明,TAM 在相同内存占用下比统一压缩的准确率更高,同时定期压缩将峰值内存限制在 3.1-3.2 GB(减少 65%),并保持具有竞争力的准确率。

英文摘要

Reasoning language models generate lengthy chain-of-thought (CoT) sequences whose key-value (KV) cache grows linearly and becomes a memory bottleneck during decoding. Existing compaction methods treat reasoning trajectories as flat token sequences and apply uniform compression, ignoring the hierarchical structure of CoT reasoning where different steps vary drastically in importance. We propose \textbf{Thought-Aware Attention Matching (TAM)}, which exploits this structure through three mechanisms: (i)~thought segmentation that decomposes the trajectory into reasoning blocks, (ii)~adaptive budget allocation that assigns compression budget based on each segment's importance and size, and (iii)~pivotal token protection that preserves high-attention reasoning anchors. We prove that the allocation rule is optimal under a convex error model and that cumulative error under sequential compaction remains bounded. Experiments on AIME 2024 and MATH-500 with Qwen3-4B show that TAM improves accuracy over uniform compaction at the same memory footprint, with periodic compaction bounding peak memory to 3.1--3.2\,GB (a 65\% reduction) while maintaining competitive accuracy.

Comments16 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑