arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35139cs.LG

CacheRepair:学习修复RAG中的跨块上下文以实现KV缓存融合

CacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache Fusion

  • The Chinese University of Hong Kong(香港中文大学)
  • Peking University(北京大学)
  • The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Genglin Wang, Wangsong Yin, Yeerzhati Abudunuer, Haoxuan Xu, Guoliang Xing, Zhenyu Yan

AI总结:

针对多文档RAG中KV缓存拼接缺乏跨块上下文的问题,提出CacheRepair轻量网络学习修复残差,在多个模型和数据集上实现质量-延迟帕累托最优,显著加速TTFT并提升F1。

AI中文摘要:

多文档检索增强生成(RAG)要求语言模型在回答问题之前处理多个检索到的文本块。独立预计算每个块的KV缓存,并在检索到这些块时拼接缓存,可以加速此步骤。然而,拼接后的缓存缺乏跨块注意力信息,降低了答案质量。选择性重计算方法通过重新运行目标LLM对选定的令牌来恢复缺失的跨块上下文,但会带来大量的在线计算。我们引入了CacheRepair,一个轻量级网络,学习独立计算的KV缓存与将块一起处理产生的KV缓存之间的差异。该网络将压缩的KV特征与令牌嵌入相结合,并使用在块内双向、从较早块流向较晚块的注意力。每个修复块接收压缩的缓存特征,并将预测的残差添加到每个文档令牌的缓存中。每个修复网络针对特定的冻结目标LLM在通用检索语料库上训练,并在下游数据集上重用。我们的分析表明,修复减少了块边界附近和块内部的KV误差。在三个目标LLM和四个下游数据集上的评估,将CacheRepair置于十二个模型-数据集组合中十一个的测量答案质量-延迟帕累托前沿上。报告的首令牌时间(TTFT)包括在线缓存传输和修复。在所有十二个组合中,最大的修复器在TTFT中位数上比完整预填充实现了1.69-4.61倍的加速,并相对于直接缓存重用将平均F1提高了2.1-26.1个百分点。

英文摘要:

Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question. Precomputing each chunk's KV cache independently and concatenating the caches when the chunks are retrieved can accelerate this step. However, the assembled cache lacks cross-chunk attention information, reducing answer quality. Selective recomputation methods recover the missing cross-chunk context by rerunning the target LLM on selected tokens, incurring substantial online computation. We introduce CacheRepair, a lightweight network that learns the difference between independently computed KV caches and those produced by processing the chunks together. The network combines compressed KV features with token embeddings and uses attention that is bidirectional within each chunk and flows from earlier to later chunks. Each repair block receives the compressed cache features, and the predicted residual is added to every document token's cache. Each repair network is trained for a specific frozen target LLM on a generic retrieval corpus and reused across downstream datasets. Our analysis shows that repair reduces KV errors both near chunk boundaries and throughout chunk interiors. Evaluation across three target LLMs and four downstream datasets places CacheRepair on the measured answer-quality-latency Pareto frontier in eleven of twelve model-dataset combinations. Reported time to first token (TTFT) includes online cache transfer and repair. Across all twelve combinations, the largest repairers achieve 1.69-4.61$\times$ speedups in median TTFT over full prefill and improve mean F1 by 2.1-26.1 percentage points over direct cache reuse.

补充信息

↑