微调一个感知KV缓存拼接的模型还是重新计算KV缓存?为何不两者兼得?
Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?
浏览论文内容
中文总结 AI 辅助
本文提出结合微调感知KV缓存拼接与选择性重算KV缓存的方法,在RULER基准上提升长上下文准确性9.7分,同时减少80%的TTFT。
中文摘要 AI 辅助
在检索增强生成(RAG)系统中,大量检索到的块被拼接起来形成输入上下文,以便用户基于外部知识获得高质量响应。因此,输入上下文长度大幅增加,导致预填充工作负载增大,进而延长了首令牌时间(TTFT)。虽然先前复用预计算键值(KV)缓存的工作有效减少了长上下文输入的TTFT,但当输入上下文变得非常长时,响应质量是否得以保持仍不清楚。在本文中,我们提出了一种组合方法,该方法(i)在考虑KV缓存拼接的情况下微调模型,并且(ii)选择性地重新计算一部分KV缓存。通过应用这两种技术,我们展示了长上下文输入准确性的提升。在RULER基准上的实验表明,对于124k令牌的输入,我们的方法相比仅重新计算KV缓存的基线,将RULER分数提高了9.7分。此外,与完全注意力相比,TTFT减少了80%。
英文摘要
In Retrieval-Augmented Generation (RAG) systems, a large number of retrieved chunks are concatenated to form the input context so that users can receive high-quality responses based on external knowledge. As a result, the input context length increases substantially, leading to a larger prefill workload and, in turn, a longer time to first token (TTFT). While previous works that reuse precomputed key-value (KV) caches effectively reduce TTFT for long-context inputs, it remains unclear whether response quality is preserved when the input context becomes very long. In this paper, we propose a combined approach that (i) fine-tunes the model while taking KV cache concatenation into account and (ii) selectively recomputes a subset of the KV caches. By applying both techniques, we demonstrate improved accuracy for long-context inputs. Experiments on the RULER benchmark show that, for a 124k-token input, our method improves the RULER score by 9.7 point over the baseline that recomputes KV caches only. Moreover, TTFT is reduced by 80% compared with full attention.
发表机构
- Kioxia Corporation(铠侠公司)
机构由 AI 辅助整理,请以论文原文为准。