AI 中文总结
该研究针对视频扩散 Transformer 自注意力计算成本高的问题,提出 Token 半径注意力机制,在多种视频生成模型配置上实现了高效加速,同时保持了良好的生成质量。
AI 中文摘要
视频扩散 Transformer(VDiT)可实现高保真生成,但密集 3D 自注意力会产生二次方计算成本。现有的头部级和块级稀疏方法在查询间共享计算预算,却忽略了每个查询特定的注意力需求。我们观察到,保留的密度随查询变化,且与注意力熵呈对数线性相关,而主导交互形成以查询为中心、具有 token 依赖半径的邻域。基于这些发现,我们提出 Token 半径注意力(TRA),这是一种无需训练的框架,可将查询熵映射为解析的 token 预算,并将其转换为时间衰减的半径,无需显式键排序。融合熵提取、预热复用和块稀疏掩码构建进一步降低了开销。在 Wan2.1、Wan2.2 和 HunyuanVideo 的 7 种文本到视频(T2V)/图像到视频(I2V)配置中,TRA 仅保留 9%-19%的注意力交互,实现 1.56 倍至 2.05 倍的加速,同时保持具有竞争力的生成质量。代码可在指定 URL 获取。
英文摘要
Video Diffusion Transformers (VDiTs) enable high-fidelity generation but incur quadratic cost from dense 3D self-attention. Existing head- and block-level sparse methods share computation budgets across queries, overlooking token-specific attention demand. We observe that retained density varies across queries yet correlates log-linearly with attention entropy, while dominant interactions form query-centered neighborhoods with token-dependent radii. Based on these findings, we propose Token Radius Attention (TRA), a training-free framework that maps query entropy to an analytic token budget and converts it into a temporally decayed radius without explicit key ranking. Fused entropy extraction, warm-up reuse, and block-sparse mask construction further reduce overhead. Across seven Wan2.1, Wan2.2, and HunyuanVideo T2V/I2V configurations, TRA retains only 9-19% of attention interactions and achieves 1.56x-2.05x speedup with competitive generation quality. Code is available at https://github.com/IF-LAB-PKU/Token-Radius-Attention.