发表机构
Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过精确分解与蒙特卡洛估计,分析KV缓存驱逐下首次分歧时机对累积不一致性的影响,发现SnapKV延迟分歧但分歧后TV仍高,分歧后暴露主导不匹配差距。
AI 中文摘要
KV缓存驱逐会扰动控制自回归生成的条件下词元分布。我们研究了首次分歧时机及随后的词元不匹配如何决定累积不一致性。我们在指定的逐步最大耦合下推导出一个精确分解:期望不匹配比例等于首次不匹配贡献加上分歧后暴露量乘以其不匹配率。在无限制的自回归核对上,一个显式构造实现了与有限分歧对齐观测窗口兼容的风险尖锐区间。残差分支条件蒙特卡洛方法提供了发生、占用及窗口/尾部贡献的无偏联合估计,并对总词元损失具有每重复方差优势。来自Meta-Llama-3.1-8B-Instruct和Qwen2.5-7B-Instruct的完整轨迹显示,在50%保留率下,SnapKV比SnapKV-512或使用相同50%提示缓存预算的近期词元保留更晚且更少地进入分歧,而分歧后总变差(TV)仍然很高。在对288篇文档的探索性分析中,分歧后暴露占四个聚合不匹配差距的85-90%。在90%保留率下对288篇独立文档,预设比较显示两个模型中后期窗口的分支对齐TV高于早期窗口。
英文摘要
KV-cache eviction perturbs the conditional token distributions governing autoregressive generation. We investigate how first-divergence timing and subsequent token mismatch determine cumulative disagreement. We derive an exact decomposition under a specified stepwise maximal coupling: the expected mismatch fraction equals a first-mismatch contribution plus post-divergence exposure multiplied by its mismatch rate. An explicit construction over unrestricted autoregressive kernel pairs realizes the sharp interval of risks compatible with a finite divergence-aligned observation window. Residual-branch conditional Monte Carlo provides unbiased joint estimates of occurrence, occupation, and window/tail contributions, with per-replicate variance dominance for total token loss. Complete trajectories from Meta-Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct show that SnapKV at 50% retention enters divergence later and less often than SnapKV-512 or recent-token retention with the same 50% prompt-cache budget, while post-divergence total variation (TV) remains high. In an exploratory analysis of 288 documents, post-divergence exposure accounts for 85-90% of four aggregate mismatch gaps. On 288 independent documents at 90% retention, prespecified comparisons show higher branch-aligned TV in the late than in the early window in both models.