arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当关键词下降但分类器保持:KV缓存压缩下的软拒绝

When Keywords Drop but Classifiers Hold: Soft Refusals under KV Cache Compression

Kang Chen, Xiuze Zhou, Hong Chen, Yuanguo Lin

arXiv 2609.31678首次发表:更新:

AI 中文总结

本研究探讨KV缓存压缩下关键词拒绝率下降而分类器保持高拒绝率的现象,提出多评判器审计方法,确保压缩后安全监控的可靠性。

AI 中文摘要

KV缓存压缩在内存受限条件下被广泛用于长上下文LLM推理,而部署系统通常在生成后通过关键词过滤器或学习分类器对拒绝进行评分。此类监控器旨在指示模型在实际使用的服务机制下是否拒绝了有害请求。然而,尚不清楚保持任务准确性的匹配压缩是否也保持了轻量级词汇监控器与更强拒绝分类器之间的一致性。我们通过配对协议对此进行研究,使用n=200个带有长填充上下文的有害提示:每个提示在完全保留和共享预填充后的匹配驱逐条件下各回答一次,并使用关键词启发式、HarmBench Llama-2-13B分类器、辅助LLM评判器以及人类对分歧进行评分。在Qwen2.5-3B上,关键词拒绝率从98.0%降至80.5%(McNemar p~1e-8),而分类器拒绝率保持在接近上限(99.0%-99.5%),MMLU准确率不变(50.0%);人类标签主要跟随分类器,与软拒绝一致。这种差距并非普遍存在,在短填充和配对SnapKV下会减弱,因此压缩下的安全审计应依赖于与服务上下文匹配的多个评判器,而非仅依赖关键词率。

英文摘要

KV cache compression is widely used for long context LLM inference under memory constraints, while deployed systems typically score refusals after generation with keyword filters or learned classifiers. Such monitors are intended to indicate whether a model declined a harmful request under the serving regime actually used. However, it remains unclear whether matched compression that preserves task accuracy also preserves agreement between lightweight lexical monitors and stronger refusal classifiers. We study this with a paired protocol on n=200 harmful prompts with a long filler context: each prompt is answered once under full retention and once under matched eviction after a shared prefill, and the same replies are scored by keyword heuristics, the HarmBench Llama-2-13B classifier, an auxiliary LLM judge, and humans on disagreements. On Qwen2.5-3B, keyword refusal falls from 98.0% to 80.5% (McNemar p~1e-8) while classifier refusal stays near ceiling (99.0%-99.5%) and MMLU accuracy is unchanged (50.0%); human labels predominantly follow the classifier, consistent with soft refusals. The gap is not universal and weakens under short fillers and paired SnapKV, so safety auditing under compression should rely on several judges matched to the serving context rather than on keyword rates alone.

Comments8 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑