arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13051cs.DScs.LG

在三注意力机制中预计算未来偏移平均值

Precomputing the Future-Offset Average in TriAttention

Amarnath Mukherjee

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对三注意力机制缩减长推理语言模型键值缓存时未来距离平均计算繁琐的问题,提出利用代数恒等式将17倍平均简化为单波段权重离线计算,使键评分成本降低,是对原方法的精确补充。

中文摘要 AI 辅助

三注意力机制是一种用于缩减长推理语言模型的键值缓存的近期方法:它通过每个缓存键可能获得的注意力对其进行评分,并淘汰得分最低的键。由于一个键不知道其未来查询的距离有多远,因此得分是在17种可能的未来距离的阶梯上进行平均的。我们指出,这种平均是不必要的:未来距离仅通过位置相关的旋转进入得分,所以整个17倍的平均精确地通过一个单行代数恒等式简化为一个单波段权重,该权重只需离线计算一次。对一个键进行评分的成本从17次评估降至1次,而被修剪的键不变。节省的成本不大,且完全体现在三注意力机制的修剪分数计算中,而非注意力内核中;我们将其作为对他们方法的一个小而精确的补充呈现,并通过数值验证了该恒等式。

英文摘要

TriAttention is a recent method for shrinking the KV cache of long-reasoning LLMs: it scores each cached key by how much attention it is likely to receive and evicts the lowest-scoring ones. Because a key does not know how far away its future queries will sit, the score is averaged over a ladder of 17 possible future distances. We point out that this average is free: the future distance enters the score only through the position-dependent rotation, so the whole 17-fold average collapses--exactly, by a one-line algebraic identity--into a single per-band weight that is computed once, offline. Scoring a key then costs one evaluation instead of seventeen, with no change to which keys get pruned. The saving is modest and lives entirely in TriAttention's pruning-score computation, not in the attention kernel; we present it as a small, exact complement to their method, and we confirm the identity numerically.

发表机构

  • Hozhoke, Inc.(霍佐克公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑