发表机构
Hangzhou Normal University; T-Head Semiconductor Co., Ltd.; Alibaba Group(杭州师范大学; 平头哥半导体有限公司; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
iCATS通过交互感知的重要性估计、SNR引导的动态稀疏性调度及尾部合并策略,在HunyuanVideo-T2V-13B和Wan2.1-T2V-14B上分别实现2.03倍、1.55倍加速,达成最优效率-质量权衡。
AI 中文摘要
无需训练的稀疏注意力通过减少计算量为扩散Transformer(DiTs)提供了实用的加速方案,且无需微调。这类方法通常需要估计查询-键区域的重要性并推导稀疏掩码,仅计算重要候选区域,这不可避免地会引入近似误差,可能降低生成质量。为更好地平衡效率与质量的权衡,本文提出iCATS,将改进的重要性估计、稀疏掩码构建与高效硬件执行策略相结合。具体而言,在重要性估计方面,与以往基于特征相似度对查询和键标记进行独立聚类以估计注意力分数的工作不同,iCATS表明基于查询-键点积交互的聚类更准确,并将该目标重新表述为简单的二次型以实现低成本计算。在稀疏掩码构建方面,iCATS未采用固定的top-p规则,而是观察到不同去噪时步对稀疏近似误差的容忍度存在差异,因此引入了SNR引导的稀疏性调度来动态调整稀疏度,从而提高准确性。最后,在硬件执行方面,iCATS设计了尾部合并策略以减少不规则聚类大小导致的填充开销,提升GPU内核利用率。大量实验表明,iCATS在HunyuanVideo-T2V-13B上实现了2.03倍加速,PSNR为31.017 dB;在Wan2.1-T2V-14B上实现了1.55倍加速,PSNR为29.301 dB,达到了当前最优的效率-质量权衡。
英文摘要
Training-free sparse attention offers a practical acceleration solution to Diffusion Transformers (DiTs) via reducing computations without fine-tuning. It typically involves estimating the importance of query-key regions and deriving sparse masks to compute only the important candidates, which inevitably introduces approximation errors that may degrade generation quality. To better balance the efficiency-quality trade-off, we propose iCATS, integrating improved importance estimation and sparse mask construction with an efficient hardware execution strategy. Specifically, for importance estimation, unlike previous works that perform independent clustering over query and key tokens based on feature similarity to estimate attention scores, iCATS demonstrates that clustering based on query-key dot-product interactions is more accurate and further reformulates this objective as a simple quadratic form for low-cost computation. For sparse mask construction, instead of using a fixed top-p rule, we observe that tolerance to sparse approximation errors varies across denoising timesteps and therefore introduce an SNR-guided sparsity schedule to adjust sparsity dynamically, leading to higher accuracy. Finally, for hardware execution, we devise a tail-merging strategy to reduce padding overhead caused by irregular cluster sizes, improving GPU kernel utilization. Extensive experiments show that iCATS achieves $2.03\times$ acceleration with 31.017 dB PSNR on HunyuanVideo-T2V-13B and $1.55\times$ acceleration with 29.301 dB PSNR on Wan2.1-T2V-14B, delivering a state-of-the-art efficiency-quality trade-off.