arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KV缓存的自适应过滤:诊断和纠正大语言模型推理中的结构角色偏差

Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference

Soumil Mandal

arXiv 2607.13205首次发表:更新:

AI 中文总结

研究大语言模型推理中KV缓存问题,通过反事实实验确定抑制KEY令牌是最佳过滤方法,采用无重新训练、基于角色的条件分配缩小与H2O方法差距,15MB线性角色探测器提供标签且推理成本可忽略。

AI 中文摘要

基于注意力的KV缓存逐出(如H2O及其衍生方法)通过根据累积注意力质量(视为信号能量)对令牌进行排序并保留最重的令牌,来压缩长上下文模型的内存受限状态。在模式密集的输入流(如嵌套JSON)上,该分数充当非平稳过滤器,不成比例地保留噪声。反事实实验表明抑制KEY令牌是最佳可部署过滤器。我们在SnapKV的窗口分数上进行无重新训练、基于角色的条件分配,由单个调优超参数控制,在低预算下缩小了H2O差距的63 - 98%,在高预算下适度匹配或超过全缓存精度。一个15MB的线性角色探测器以可忽略的推理成本提供这些标签。

英文摘要

Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking tokens on accumulated attention mass, treated here as signal energy, and keeping the heaviest. On schema-dense input streams such as nested JSON, this score acts as a non-stationary filter that disproportionately retains noise: a non-content sink role (delimiters or whitespace) carries an order of magnitude more energy than any content role, and structural KEY tokens are over-retained at roughly 1.8x the rate of the answer-carrying VALUE tokens, collapsing exact-match accuracy from 88% to 0% at a 5% budget as the signal-to-noise ratio of the retained state degrades. A counterfactual experiment establishes that suppressing KEY tokens is the best deployable filter. Our retraining-free, role-conditional allocation over SnapKV's windowed score, governed by a single tuned hyperparameter, closes 63-98% of the H2O gap at sub-20% budgets and, at higher budgets, modestly matches or exceeds full-cache accuracy -- a small, seed-sensitive denoising effect (borderline significant at B=0.50; not distinguishable from zero at B=0.30 over four seeds). A 15 MB linear role probe supplies these labels at negligible inference cost, though matching parser-level downstream accuracy remains open.

Comments6 pages, 2 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑