arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26965cs.LGcs.CV

ClusterAttention:一种无需训练的双向注意力加速方法

ClusterAttention: A training-free speedup of bidirectional attention

Kasper Nordenram, Amelie Dittmann

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出无需训练的ClusterAttention,通过适配键查询几何的快速递归聚类实现双向注意力加速,在表格数据上提速2-6倍、精度保留≥99%,在视频生成上比SVOO加速效果更好。

中文摘要 AI 辅助

本文提出了ClusterAttention,这是一种通用的无需训练的双向注意力层加速方法。现有的稀疏注意力方法要么依赖输入中的结构,例如语言中的顺序或图像中的空间邻近性,要么使用在多次前向传播中摊销的慢速聚类过程。ClusterAttention则采用快速递归聚类方法,该方法会适配每个注意力头中键(key)和查询(query)的几何结构,以生成有用的聚类。该方法允许任意设置聚类的大小,我们通过将所有聚类设置为2的幂次的固定大小,使基于GPU的块稀疏注意力在每对查询-键交互上的延迟与密集注意力相同。我们还推导了稀疏注意力的输出误差表达式,该表达式解释了一个反直觉的实验发现:紧密聚类比随机聚类会导致更大的误差。随后,我们推导了通过聚类中心补偿被排除聚类时的误差,并表明该误差会随聚类变紧密而缩小。我们将这种补偿机制集成到方法中。在大规模表格数据上,ClusterAttention使TabPFN-3 arXiv:2605.13986的运行速度提高了2至6倍,同时保留了至少99%的密集注意力精度。据我们所知,这是第一种可成功应用于非结构化输入和单次前向传播场景的无需训练的方法。在使用Wan 2.1-14B T2V arXiv:2503.20314进行视频生成时,ClusterAttention生成的输出比密集注意力更接近真实结果,且与专门针对该领域开发的领先方法SVOO arXiv:2603.18636相比,其加速比更大(1.8倍对比1.4倍),且两者均无需离线校准即可运行。

英文摘要

We introduce ClusterAttention, a general training-free speedup of bidirectional attention at large token counts. We point out two common assumptions in contemporary training-free methods; attention sparsity, and context that can be leveraged, such as structure in the input or multiple similar forward passes, and show when they fail. Our proposed method utilizes a fast attention-aware recursive clustering method, and compensation of excluded clusters through their mean. The clustering method gives power-of-two cluster sizes, allowing block-sparse attention to match dense attention in GPU throughput. On TabPFN-3 arXiv:2605.13986, a model where none of the assumptions hold, ClusterAttention is to our knowledge the first method to provide a substantial speedup over the default attention, while consistently keeping over 99\% of its accuracy. On the largest dataset from the TALENT benchmark suite, it makes processing of the training dataset close to 8x faster at nearly 11x attention speedup. ClusterAttention is also competitive with domain-specific methods, while avoiding any of the domain-specific engineering. On video-generation with Wan 2.1-T2V-14B arXiv:2503.20314 it produces output closer to dense attention at a larger speedup (1.8x vs 1.4x) than SVOO arXiv:2603.18636, a leading method in this domain, with both evaluated without offline calibration.

补充信息

↑