KV-Rescue:通过逐步交错恢复推理语言模型的KV驱逐损失
KV-Rescue: Recovering Reasoning Language Model KV Eviction Loss via Stepwise Interleaving
AI总结:
KV-Rescue是一种无需训练的推理框架,通过交错轻量级辅助模型的推理步骤,结合在线检测器,在驱逐预算B=64时平均恢复了87%的KV驱逐损失准确率,还减少了43%的基础模型token生成量。
AI中文摘要:
KV缓存驱逐限制了长推理轨迹的内存成本,但本质上是有损的,因为模型仅能从部分历史信息中解码。在严格的预算下,这不仅会降低准确率,还可能引发失控退化,即模型会生成不连贯或重复的token,直至达到长度限制。我们将大部分此类损失归因于缺失上下文导致的信息缺口,而非模型容量有限导致的能力缺口。被驱逐的7B模型与完整上下文的1.5B模型会产生互补错误,而对它们的答案进行最优选择可恢复完整KV 7B模型准确率差距的79%。基于此观察,我们提出KV-Rescue,这是一种无需训练的推理框架,通过轻量级完整上下文辅助模型弥合KV驱逐带来的信息缺口。KV-Rescue将两个模型的推理步骤交错到共享轨迹中。在线检测器利用熵和可压缩性提前终止不连贯或重复的基础模型候选的生成。在使用Qwen2.5-Math 7B和72B的五个数学基准测试中,当驱逐预算B=64时,KV-Rescue平均恢复了87%因驱逐损失的准确率。解码成本分析进一步显示,防止失控退化平均减少了基础模型43%的token生成量。
英文摘要:
KV-cache eviction caps the memory cost of long reasoning traces but is inherently lossy because the model decodes from a partial view of its history. Under aggressive budgets, this not only lowers accuracy but can also cause runaway degeneration, where the model produces incoherent or repetitive tokens until reaching the length limit. We characterize much of this loss as an information gapf caused by missing context, rather than a capability gap caused by limited model capacity. An evicted 7B model and a full-context 1.5B model make complementary errors, and an oracle choice between their answers recovers 79% of the accuracy gap to the full-KV 7B model. Based on this observation, we propose KV-Rescue, a training-free inference framework that bridges the information gap introduced by KV eviction using a lightweight full-context helper. KV-Rescue interleaves reasoning steps from the two models into a shared trajectory. An online detector uses entropy and compressibility to terminate the generation of incoherent or repetitive base-model candidates early. Across five math benchmarks with Qwen2.5-Math 7B and 72B, KV-Rescue recovers an average of 87% of the accuracy lost to eviction at eviction budget B=64. A decode-cost analysis further shows that preventing runaway degeneration cuts base-model token generation by 43% on average.