arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OVAL:面向输出的局部页基用于KV缓存检索

OVAL: Output-Aware Local Page Bases for KV Cache Retrieval

Ashkan Shahbazi, Chayne Thrash, Soheil Kolouri

arXiv 2610.06686首次发表:更新:

发表机构

College of Connected Computing; Vanderbilt University(互联计算学院; 范德比尔特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长上下文推理中KV缓存检索,提出面向输出的页编码OVAL,利用键值联合结构提升检索准确性,无需训练,在多个基准上优于仅键基线。

AI 中文摘要

随着注意力机制必须处理不断增长的KV缓存,大型语言模型的长上下文推理变得越来越昂贵。页稀疏注意力通过紧凑地表示每个KV页并仅为每个查询检索子集来降低这一成本。现有的检索方法旨在估计注意力分数或页面相关性,但其目标并未直接考虑近似误差如何影响由此产生的值加权注意力输出。我们引入了\new{},一种面向输出的页面编码,它源自键和值的联合结构,同时保留了准确检索所需的关键信息。\new{}无需训练,在推理时不需要额外的值相关统计量。一旦构建完成,其存储表示的大小和解码时间评分成本与仅键的频谱表示相同。在长推理、长上下文理解和长生成基准测试中,\new{}持续优于仅键的频谱基线,并与最近的KV缓存压缩和检索方法表现相当。在长推理基准测试中,它在模型-基准对上实现了强大的avg@k性能,同时在多个长上下文理解和生成设置中匹配或超越领先的基线,且解码开销适中。代码可在\new{}获取。

英文摘要

Long context inference with large language models becomes increasingly expensive as attention must operate over an ever growing KV cache. Page sparse attention reduces this cost by representing each KV page compactly and retrieving only a subset for each query. Existing retrieval methods are designed to estimate attention scores or page relevance, but their objectives do not directly account for how approximation errors affect the resulting value weighted attention output. We introduce \method{}, an output aware page encoding derived from the joint structure of keys and values while preserving the key information needed for accurate retrieval. \method{} is training free and requires no additional value dependent statistics at inference time. Once constructed, its stored representation has the same size and decode time scoring cost as a key only spectral representation. Across long reasoning, long context understanding, and long generation benchmarks, \method{} consistently improves over the key only spectral baseline and performs competitively with recent KV cache compression and retrieval methods. On long reasoning benchmarks, it achieves strong avg@\(k\) performance across model benchmark pairs, while matching or surpassing leading baselines on several long context understanding and generation settings with modest decoding overhead. Code is available at \url{https://github.com/Ashkan13776/oval-kv}.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑