超越稀疏权重:注意力机制何时可压缩?
Beyond Sparse Weights: When Is Attention Compressible?
浏览论文内容
中文总结 AI 辅助
该研究针对KV缓存压缩的假设缺陷,提出无需训练的压缩器CertKV,在多个长上下文基准测试中表现优异,明确注意力可压缩性的决定因素。
中文摘要 AI 辅助
KV缓存压缩常基于注意力图中存在少量大权重这一假设,但该假设并不完整:大权重可能不包含大部分质量,被省略的值会相互抵消,保留注意力输出未必能保留任务性能。我们将这些问题分开研究:全局分数差距而非阈值计数决定了保留目标质量所需的token数量;对于已实现的行,被省略值的加权和是精确的缺失统计量;受控检索-聚合模型解释了截断何时有益、何时有害。这些结果催生了CertKV——一种无需训练的压缩器,它为每个注意力头预留一个尾部摘要槽,其余槽按值的离散度分配。在匹配预算下,CertKV在9项LongBench-v2设置中7项排名前二,在128K RULER上仍处于领先压缩层级,在打包的Llama原型中实现了10倍的缓存预算。可压缩性取决于质量、值、未来查询和任务,而非仅取决于看似稀疏的注意力图。
英文摘要
KV-cache compression is often justified by attention maps with a few large weights. This is incomplete: large weights may not contain most of the mass, omitted values can cancel, and preserving the attention output may not preserve the task. We separate these questions. Global score gaps -- not threshold counts -- determine how many tokens are needed to retain a target mass. For a realized row, the weighted sum of omitted values is the exact missing statistic. A controlled retrieval--aggregation model explains when truncation helps and when it hurts. These results motivate CertKV, a training-free compressor that reserves one tail-summary slot per head and allocates the rest by value dispersion. Under matched budgets, CertKV is top-two in seven of nine LongBench-v2 settings, remains in the leading compressed tier on 128K RULER, and realizes a ten-fold cache budget in a packed Llama prototype. Compressibility depends on the mass, values, future queries, and task -- not on a sparse-looking map alone.
发表机构
- City University of Hong Kong(香港城市大学)
- University of New South Wales(新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。