Cartridges++:KV缓存压缩,避免脱离上下文时的性能滑坡
Cartridges++: KV Cache Compression without Off-Context Derailment
- Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出Cartridges++,通过路由器或数据混合的简单修改,在保持KV缓存压缩性能的同时,恢复模型处理上下文外查询的能力,避免仅评估文档效用而忽视模型整体能力退化。
AI中文摘要:
向大型语言模型(LLM)重复提供长文档成本高昂:计算量随上下文长度增长,键值(KV)缓存的内存占用也会急剧膨胀。压缩KV(CKV)表示旨在模拟文档的缓存,通常在一次计算后即可在推理前永久使用。获取CKV的方法从减少列数的丢弃机制到学习式方法不等。在学习式方法中,Cartridges已成为领先的压缩方法,通过在相关问答对上进行蒸馏来学习紧凑的KV表示。现有评估主要关注Cartridges和其他CKV能否在文档相关的上下文内查询中产生大致相似的响应,而我们则研究关键的部署问题:它们能否处理上下文外查询,而这正是原生KV表示凭借注意力机制尤为擅长的。我们观察到一种基本权衡:Cartridges在上下文内查询上表现更好,而启发式变体则更好地保留了原始LLM在上下文外操作的能力。我们通过它们避免响应中上下文污染、保留通用知识以及遵循指令的能力来衡量这一点。我们提出Cartridges++,这是对Cartridges的简单修改,以微小或可忽略的代价保留上下文外能力。路由器变体在推理时决定查询是否应使用学习到的长上下文记忆,而数据混合变体则将一小部分训练问答分配给参考长文档之外的查询。我们的研究表明,仅凭文档效用评估CKV可能会掩盖更广泛模型能力的显著退化,但这些问题可以通过对CKV推理或训练进行良性修改来解决。
英文摘要:
Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-value (KV) cache balloons. Compressed KV (CKV) representations aim to mimic the cache of a document and are typically computed once and for all, ahead of inference time. Methods to obtain CKVs range from drop mechanisms that reduce their number of columns, to learned approaches. Among the latter, Cartridges have emerged as a leading compression method, learning compact KV representations through distillation on relevant Q/A pairs. While existing evaluations focus primarily on whether Cartridges and other CKVs yield approximately similar responses to document-related, on-context queries, we investigate the crucial deployment question of whether they can handle off-context queries, something the native KV representation is particularly good at, thanks to the mechanics of attention. We observe a fundamental trade-off: while Cartridges perform better for on-context queries, heuristic-variants preserve better the original LLM's ability to operate off-context. We measure this through their capability to avoid context contamination in their response, retain general knowledge, and follow instructions. We propose Cartridges++, simple modifications to cartridges that retain off-context abilities at small or negligible cost. The router variant decides at inference time whether the query should use the learned long-context memory, while the data-mixing variant allocates a small fraction of training Q/As to queries outside the reference long document. Our study shows that assessing CKVs on document utility alone can mask substantial degradation in broader model capabilities, yet those issues can be fixed with benign changes to CKV inference or training.