arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ATTUNER:通过查询侧适配实现免重计算的KV缓存复用

ATTUNER: Recomputation-Free KV Cache Reuse via Query-Side Adaptation

Xinghao Chen, Junnan Dong, Cai Ke, Chak Tou Leong, Haocheng Sun, Keyu Chen, Siyu An, Ruizhi Qiao, Xing Sun, Wenjie Li, Xiaoyu Shen

arXiv 2609.36722首次发表:更新:

发表机构

Eastern Institute of Technology; The Hong Kong Polytechnic University; Tencent Youtu Lab(东方理工大学; 香港理工大学; 腾讯优图实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ATTUNER通过查询侧低秩适配器学习读取冻结的工件缓存,无需重计算KV即可复用,在多个基准上匹配全上下文预填充质量并实现3.73倍加速。

AI 中文摘要

大型语言模型(LLM)智能体在将技能、文档和记忆条目等可复用内容加载到当前上下文时,会反复进行编码。每次请求都重新编码这些内容会浪费计算资源。位置无关缓存(PIC)通过独立编码每个工件并在任意位置复用其键值(KV)状态来缓解这一问题,但与全上下文预填充相比,它会导致质量损失。现有方法通过恢复全局位置ID或重计算选定token来修复这一损失。在本工作中,我们隔离了损失的来源,发现位置不匹配的影响较小,独立缓存的工件保留了忠实的表示:读取提供的工件在很大程度上保持准确,只有当模型需要在多个工件之间进行选择时,性能才会下降。此外,用全预填充分数替换PIC的注意力分数可以在缓存KV不变的情况下恢复性能,将失败定位在注意力而非KV重计算上。受此启发,我们提出了ATTUNER,一种查询侧适配方法,学习读取冻结的工件缓存。ATTUNER在查询投影中插入低秩适配器,并通过将全预填充分布蒸馏到学生模型中进行训练。它训练的模型参数少于0.05%,并且在推理时既不需要缓存重计算,也不需要全上下文参考。在Qwen3-4B和Qwen3-8B上,跨越涵盖技能、文档、记忆和代码的七个基准测试中,ATTUNER在域内和域外设置中均大幅优于先前的PIC基线,匹配全上下文预填充质量,同时提供高达3.73倍的加速。

英文摘要

Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching (PIC) alleviates this by encoding each artifact independently and reusing its key-value (KV) states at arbitrary positions, but it incurs a quality loss relative to full-context prefill. Existing methods repair this loss by restoring global position IDs or recomputing selected tokens. In this work, we isolate the source of the loss, finding that the positional mismatch has minor effect, and independently cached artifacts retain faithful representations: reading a provided artifact stays largely accurate, and performance degrades only when the model must select among multiple artifacts. Moreover, replacing PIC's attention scores with full-prefill scores recovers performance with the cached KV unchanged, localizing the failure to the attention rather than KV recomputation. Motivated by this, we propose \textsc{Attuner}, a query-side adaptation method that learns to read a frozen artifact cache. \textsc{Attuner} inserts low-rank adapters into the query projections and is trained by distilling full-prefill distribution into the student. It trains fewer than 0.05\% of the model parameters and, at inference, requires neither cache recomputation nor a full-context reference. On Qwen3-4B and Qwen3-8B across seven benchmarks covering skills, documents, memory, and code, \textsc{Attuner} substantially outperforms prior PIC baselines in both in-domain and out-of-domain settings, matches full-context prefill quality while providing up to $3.73\times$ speedup.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑