准确率等同于证据吗?KV缓存压缩下的推理忠实性
Does Accuracy Equal Evidence? Reasoning Faithfulness under KV Cache Compression
浏览论文内容
中文总结 AI 辅助
该研究针对大型推理模型,发现KV缓存压缩中存在答案-证据差距,即令牌驱逐类方法可保留高准确率却大幅降低推理忠实性,而量化方法受影响更小,表明失效源于推理轨迹部分丢失。
中文摘要 AI 辅助
KV缓存压缩通常通过最终答案准确率进行评估,隐含假设是保留答案也会保留支撑该答案的推理过程。我们针对大型推理模型检验这一假设,发现其可能不成立:在压缩下,正确答案与其可见支撑推理依据的有效性会以不同的比例被保留。我们采用受控固定轨迹重放协议研究该失效问题,该协议固定推理内容,隔离压缩是否保留已可用轨迹中的可用信息。我们在三个模型上针对数学推理、科学问答、临床计算和长上下文检索任务,评估了十种令牌驱逐类KV压缩方法和一种量化方法。我们测量了最终准确率、答案链一致性和扰动忠实性。在所有任务中,令牌驱逐方法可保留有竞争力的最终答案准确率,却会大幅降低链支撑或扰动忠实性,我们将此称为“答案-证据差距”。作为对照的保留覆盖范围的量化方法受影响显著更小,表明该失效与KV内存减少本身关联较小,而与丢失部分推理轨迹的访问权关联更大。代码可在该https URL获取。
英文摘要
KV cache compression is commonly evaluated by final-answer accuracy, implicitly assuming that preserving the answer also preserves the reasoning that supports it. We test this assumption for large reasoning models and show that it can fail: under compression, correct answers and the validity of their visible supporting rationales can be preserved at different rates. We study this failure with a controlled fixed-trace replay protocol, which holds reasoning content fixed and isolates whether compression preserves usable information from an already available trace. We evaluate ten token-eviction KV compression methods and one quantization method on three models across mathematical reasoning, scientific QA, clinical calculation, and long-context retrieval. We measure final accuracy, answer-chain consistency, and perturbation faithfulness. Across tasks, token-eviction methods can preserve competitive final-answer accuracy while substantially degrading chain support or perturbation faithfulness. We call this the answer-evidence gap. A coverage-preserving quantization control is substantially less affected, suggesting that the failure is tied less to KV memory reduction itself than to losing access to parts of the reasoning trace. Code is available at https://github.com/famous-blue-raincoat/Safe_KV_Compress.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。