arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KVDiagnosis:长上下文语言模型中KV缓存压缩的诊断基准

KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models

Chen Qiu, Ziwu Liu, Chao Fei, Guozhong Li, Panos Kalnis

arXiv 2608.09412首次发表:更新:

发表机构

KAUST(阿卜杜拉国王科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出KVDiagnosis,即长上下文语言模型KV缓存压缩的诊断基准,通过分类方法、设计评估流程、建立记录格式,在Qwen3-8B上验证其可有效区分压缩成败,代码与数据公开。

AI 中文摘要

KV缓存压缩可减少长上下文内存消耗,但聚合任务分数无法揭示哪些正确执行会失败以及失败原因。本文提出KVDiagnosis,这是一个诊断数据集和基准,包含三项贡献:第一,25种方法的分类将方法归为5个机制家族,并将其与8个经过验证的实现及其有效诊断测量关联;第二,对于每个支持的方法设置,我们在每个固定拆分中针对每个源,以每个源的FullCache(完整缓存)作为对照,分别选择FullCache正确/压缩错误(C-to-W)行,因此没有压缩器定义另一个的测试集;第三,通用记录格式将配对输出和运行元数据与缓存、似然、注意力和解码测量关联,并带有明确的适用性状态。在Qwen3-8B上,4种证据感知工作负载在2600个源上产生59800个支持的压缩运行,以及12520个C-to-W行。在固定诊断规则下,63.2%的运行具有低或部分测量/预测覆盖率;仅19行(0.2%)结合了高测量/预测覆盖率与强似然漂移;另有2126行(17.0%)保留了结构位置可寻址性,其表示保真度仍未知,同时表现出相同的漂移。与C-to-C(完整缓存正确)成功对照相比,所有10种诊断均能区分失败与成功的压缩(分层AUROC为0.684-0.871)。在96个可复现的低EAR(证据感知率)失败中,受控的4倍证据注意力提升修复了29.2%,而在计数匹配的虚假干预下修复率为6.3%,在匹配的C-to-C对照上则出现3.3%的性能下降。代码和数据可在该https URL获取。

英文摘要

KV-cache compression reduces long-context memory, but aggregate task scores reveal neither which correct executions fail nor why. We present KVDiagnosis, a diagnostic dataset and benchmark with three contributions. First, a 25-method taxonomy groups methods into five mechanism families and links them to eight verified implementations and their valid diagnostic measurements. Second, for every supported method setting, we evaluate all sources in each fixed split against a per-source FullCache control before selecting FullCache-correct/compressed-wrong (C-to-W) rows separately for each method-setting, so no compressor defines another's test set. Third, a common record format links paired outputs and run metadata to cache, likelihood, attention, and decoding measurements with explicit applicability states. On Qwen3-8B, four evidence-aware workloads yield 59 800 supported compressed runs over 2600 sources and 12 520 C-to-W rows. Under fixed diagnostic rules, 63.2% have low or partial measured/projected coverage. Only 19 rows (0.2%) combine high measured/projected coverage with strong likelihood drift; another 2,126 (17.0%) preserve structural position addressability, for which representation fidelity remains unknown, while showing the same drift. Against C-to-C success controls, all ten diagnostics separate failed from successful compression (stratified AUROC 0.684-0.871). Among 96 reproducible low-EAR failures, a controlled 4x evidence-attention boost repairs 29.2%, versus 6.3% under a count-matched sham intervention and 3.3% degradation on matched C-to-C controls. Code and data are available at https://github.com/ChosenQC/KVDiagnosis.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑