arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03228cs.LGstat.AP

SAKI:面向长上下文KV检索的分数感知低秩键索引

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval

Lin Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出无需训练的KV缓存索引SAKI,通过优化注意力分数目标,在多个开源大模型上的长上下文KV检索任务中,相比键PCA方法显著提升了前k个召回率。

中文摘要 AI 辅助

现有低秩KV缓存方法要么保留模型权重,要么保留键方差,均未直接反映推理过程中使用的注意力分数。我们推导了由秩r键压缩导致的期望注意力分数失真,并表明该失真会产生协方差加权低秩目标。在边际条件下,控制此失真还可提升前k个召回率。最优秩r解具有闭式非对称分解形式,可通过协方差加权键查询算子的奇异值分解(SVD)获得。这催生了SAKI,一种无需训练的KV缓存索引,它直接保留注意力分数而非键重构质量。在LLaMA 3.1 8B、Qwen 2.5 7B、Mistral 7B v0.1和Llama 3.2 3B模型上,SAKI在所有测试的秩下均优于键PCA方法。在秩32时,它消除了PCA剩余前64个召回误差的13%至30%,包括LLaMA 3.1 8B上从0.748提升至0.799、Qwen 2.5 7B上从0.786提升至0.850的改进。它在每个模型中提升了68%至89%的注意力头,最深层的提升最大。预测分数均方误差(MSE)降低与经验测量结果高度匹配,皮尔逊相关系数达0.997;而消融研究证实,这些改进源于优化注意力分数目标,而非仅协方差加权。对评分算子的分析进一步解释了为何仅权重、不变子空间和键重构方法可能并非最优。

英文摘要

Existing low rank KV cache methods preserve either model weights or key variance, neither of which directly reflects the attention scores used during inference. We derive the expected attention score distortion caused by rank r key compression and show that it yields a covariance weighted low rank objective. Under a margin condition, controlling this distortion also improves top k recall. The optimal rank r solution has a closed form asymmetric factorization obtained from the SVD of the covariance weighted query key operator. This motivates SAKI, a training free KV cache index that directly preserves attention scores rather than key reconstruction quality. Across LLaMA 3.1 8B, Qwen 2.5 7B, Mistral 7B v0.1, and Llama 3.2 3B, SAKI outperforms key PCA at every tested rank. At rank 32, it removes 13 to 30 percent of PCA's remaining top 64 recall error, including improvements from 0.748 to 0.799 on LLaMA 3.1 8B and from 0.786 to 0.850 on Qwen 2.5 7B. It improves 68 to 89 percent of attention heads per model, with the largest gains in deeper layers. Predicted score MSE reductions closely match empirical measurements, with a Pearson correlation of 0.997, while ablation studies confirm that the gains arise from optimizing the attention score objective rather than covariance weighting alone. Analysis of the scoring operator further explains why weight only, invariant subspace, and key reconstruction methods can be suboptimal. SAKI uses random-matrix theory to separate genuine covariance signal from autocorrelated sampling noise, matching PCA with only 512 calibration tokens and adding value exactly where PCA sees no reliable signal.

补充信息

↑