感知压缩的弃权:当KV压缩掩码移除答案证据时,教大型语言模型拒绝回答
Compression-Aware Abstention: Teaching LLMs to Refuse When KV-Compression Masks Remove Answer Evidence
浏览论文内容
中文总结 AI 辅助
该研究首次将感知压缩的弃权表述为学习问题,通过训练LoRA适配器,在减少LLM因KV缓存压缩产生的幻觉的同时保留正确回答能力,在压缩缓存解码下取得显著提升。
中文摘要 AI 辅助
KV缓存压缩通过驱逐上下文令牌来减少大型语言模型(LLM)的推理内存,但当被驱逐的令牌包含承载答案的证据时,模型可能会产生幻觉,而非意识到压缩后的上下文信息不足。我们从行为视角解决这一失败问题:据我们所知,这是首个将感知压缩的弃权(abstention)表述为学习问题的研究,其中模型学习在支持证据经压缩后保留时回答,在证据被移除时弃权(不执行)。我们从压缩器存活掩码和紧密的答案承载跨度中构建监督信号,当证据存活时将示例标记为“置信”,当证据被移除时标记为“弃权”。在约2600个MuSiQue 2跳问答示例上训练的1010万参数LoRA适配器,在提示式截断下将基础模型的幻觉减少了97%,同时在证据保留的示例上保留了正确回答能力。与仅提示的弃权基线(其在许多可回答的高保留示例上过度弃权)不同,训练后的适配器学习了一种条件策略。我们还在实际压缩缓存解码下评估了该方法,其中多压缩器训练在证据保留示例上相对于未辅助的基础模型实现了6至22倍的相对提升。受控删除实验表明,学习到的行为仅由证据内容而非输入长度驱动。
英文摘要
KV-cache compression reduces LLM inference memory by evicting context tokens, but when the evicted tokens contain answer-bearing evidence, the model may hallucinate instead of recognizing that the compressed context is insufficient. We address this failure from a behavioral perspective: to our knowledge, this is the first work to formulate compression-aware abstention as a learning problem, in which a model learns to answer when supporting evidence survives compression and abstain when it does not. We construct supervision from compressor survival masks and tight answer-bearing spans, labeling examples as Confident when evidence survives and Abstain when it is removed. A 10.1M-parameter LoRA adapter trained on ~2.6K MuSiQue 2-hop QA examples reduces base-model hallucinations by 97% under prompt-style truncation while preserving correct answering on evidence-retaining examples. Unlike prompt-only abstention baselines, which over-abstain on many answerable high-retention examples, the trained adapter learns a conditional policy. We also evaluate the method under actual compressed-cache decoding, where multi-compressor training yields a 6-22x relative lift over the unaided base on evidence-retaining examples. Controlled-deletion experiments show that the learned behavior is driven by evidence content rather than input length alone.
发表机构
- University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。