PARK:视频扩散Transformer中稀疏注意力的精确块检索
PARK: Accurate Block Retrieval for Sparse Attention in Video Diffusion Transformers
查看机构详情
- South China University of Technology(华南理工大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
PARK提出一种无需训练的稀疏注意力方法,通过保留原始查询并利用查询信息变换键来解决块检索失配,在视频扩散Transformer中实现精确块检索,提升生成质量并加速推理。
中文摘要 AI 辅助
扩散Transformer(DiTs)已成为视频生成的主导架构,但其效率受限于全注意力的二次复杂度。稀疏注意力通过检索重要块并仅在这些块内计算注意力来降低这一成本,但不准确的检索要么降低生成质量,要么产生不必要的计算。我们识别出在使用查询块和键块的平均表示进行块检索的方法中存在的两种检索失配:(i)查询侧聚合失配,即在Softmax之前对查询进行平均无法保留它们各自的注意力偏好;(ii)键侧聚类度量失配,即在原始键空间中使用标准欧氏聚类可能会在当前查询下将具有不同QK分数的键分组在一起,因此它们的平均表示可能无法准确表示当前查询如何对各个键打分。这些失配可能导致不准确的块检索。为解决这些失配,我们提出PARK,一种无需训练的稀疏注意力方法,用于精确块检索。PARK保留每个原始查询,独立地对其在键块上的注意力进行归一化,然后在每个查询块内平均这些分布。它还利用当前查询的信息在聚类前变换键,从而使获得相似QK分数的键被分组在一起。一个融合GPU内核进一步减少了块检索的开销。在HunyuanVideo和Wan上的实验表明,PARK提高了块检索准确性并保持了生成质量,同时加速了推理,在所比较的稀疏注意力方法中实现了最佳的质量-效率权衡。
英文摘要
Diffusion Transformers (DiTs) have become a dominant architecture for video generation, but their efficiency is limited by the quadratic complexity of full attention. Sparse attention reduces this cost by retrieving important blocks and computing attention only within them, but inaccurate retrieval can either degrade generation quality or yield unnecessary computation. We identify two retrieval mismatches in methods that retrieve blocks using the averaged representations of query and key blocks: (i) query-side aggregation mismatch, where averaging queries before Softmax fails to preserve their individual attention preferences, and (ii) key-side clustering metric mismatch, where standard Euclidean clustering in the original key space can group keys with dissimilar QK scores under the current query, so their average representation may not accurately represent how the current query scores individual keys. These mismatches can lead to inaccurate block retrieval. To address these mismatches, we propose PARK, a training-free sparse attention method for accurate block retrieval. PARK retains every original query, independently normalizes its attention over key blocks, and then averages these distributions within each query block. It also uses information from the current queries to transform keys before clustering, so that keys receiving similar QK scores are grouped together. A fused GPU kernel further reduces the overhead of block retrieval. Experiments on HunyuanVideo and Wan demonstrate that PARK improves block retrieval accuracy and preserves generation quality while accelerating inference, achieving the best quality-efficiency trade-off among the compared sparse attention methods.