发表机构
Air University; National University of Sciences and Technology (NUST); University of Oklahoma; Lahore University of Management Sciences (LUMS)(空军大学; 国家科学技术大学; 俄克拉荷马大学; 拉合尔管理科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对检索增强生成的KV缓存量化,以Qwen2.5-7B-Instruct为对象,发现INT4量化会严重损害忠实性且准确性指标无法察觉,强调需在部署压缩缓存前审计忠实性。
AI 中文摘要
检索增强生成系统可预计算并存储检索文档的键值(KV)缓存,以避免每次查询时重新编码上下文。对这些缓存进行量化可进一步减少存储,但此前尚无研究探讨压缩是否会损害忠实性,即响应是否仍基于检索到的证据。忠实性与准确性并不等价:模型可能生成正确答案,但该答案不再受给定上下文支持。我们在RGB和HotpotQA数据集上评估Qwen2.5-7B-Instruct模型的INT8和INT4量化效果,借助幻觉检测器、自然语言推理(NLI)蕴含关系及大语言模型(LLM)评判器,同时测量准确性与忠实性。INT8在两项指标上均接近无损;INT4会降低准确性,更关键的是,在事实正确的答案中,超过90%的忠实性变化为负面,即准确性指标对这种下降视而不见。在检索存在噪声或检索块数量增多时,损害会加剧。压缩缓存部署前必须对忠实性进行审计。
英文摘要
Retrieval-augmented generation systems can precompute and store key-value caches of retrieved documents to avoid re-encoding context at every query. Quantizing these caches further reduces storage, but no prior work asks whether compression damages faithfulness, whether responses remain grounded in the retrieved evidence. Faithfulness and accuracy are not equivalent: a model can produce a correct answer that is no longer supported by the context it was given. We evaluate Qwen2.5-7B-Instruct under INT8 and INT4 quantization on RGB and HotpotQA, measuring both accuracy and faithfulness with a hallucination detector, NLI entailment, and an LLM judge. INT8 is near-lossless across both metrics. INT4 reduces accuracy and, more critically, even among answers that remain factually correct, over 90% of faithfulness changes are negative, i.e., accuracy metrics are blind to this regression. The harm grows under noisy retrieval and with more retrieved chunks. Faithfulness must be audited before compressed caches are deployed.
CommentsGrounding Language Models: Learning Faithfully and Efficiently @ EMNLP 2026