arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12032cs.CVcs.AI

LoSA:用于无训练视频扩散加速的近无损稀疏注意力

LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration

Enhuai Liu, Yunke Wang, Yutong Wang, Changming Sun, Chang Xu

首次发表
浏览论文内容

中文总结 AI 辅助

LoSA是一种无训练稀疏注意力方法,通过固定99%的注意力质量阈值移除冗余计算,在多款视频扩散Transformer上实现最高3.2倍加速,且速度-质量权衡优于现有基线。

中文摘要 AI 辅助

视频扩散Transformer采样成本高昂:每一步去噪都会在长3D token序列上应用自注意力,随着分辨率和时长增加,二次计算成本会成为主导。稀疏注意力可在无需重新训练的情况下降低该成本,但现有方法追求激进的稀疏性,进一步提速时会不成比例地损失注意力保真度。我们聚焦于该权衡的另一端:通过构造固定近无损的保真度,并在该约束允许的范围内尽可能移除计算量。两项观察使该方案可行:约40%的块交互可被移除,同时保留99%的注意力质量,且高质量支持在去噪步骤间保持稳定。我们提出LoSA,一种无训练的稀疏注意力方法,它固定保留质量阈值为99%而非稀疏率:在一个早期密集步骤测量精确的块注意力质量,对每个头和查询块,保留满足阈值的最小键/值块集,并将冻结的块索引复用至所有后续步骤。在Wan2.1-1.3B上,仅LoSA即可实现1.36倍加速,VBench整体指标下降0.06点。该优势在组合场景下最为显著:与特征缓存结合时,LoSA在HunyuanVideo上实现3.2倍加速,指标下降0.02点,而在可比速度下,最强的稀疏基线指标下降0.32点。在三个视频扩散Transformer上,加速最高达3.2倍,LoSA始终实现最佳的无训练速度-质量权衡。

英文摘要

Video diffusion transformers are costly to sample: every denoising step applies self-attention over a long 3D token sequence, a quadratic cost that dominates as resolution and duration grow. Sparse attention reduces this cost without retraining, but existing methods pursue aggressive sparsity, where further speedup costs disproportionately more attention fidelity. We target the opposite end of this trade-off: fix near-lossless fidelity by construction, and remove as much computation as this constraint permits. Two observations make this regime practical: roughly 40% of block interactions can be removed while retaining 99% of the attention mass, and the high-mass support remains stable across denoising steps. We propose LoSA, a training-free sparse-attention method that fixes a retained-mass threshold of 99% rather than a sparsity ratio: it measures exact block attention masses at one early dense step, keeps, for each head and query block, the smallest key/value block set meeting the threshold, and reuses the frozen block indices for all remaining steps. On Wan2.1-1.3B, LoSA alone gives a $1.36\times$ speedup with a 0.06-point VBench Overall drop. The benefit is largest under composition: combined with feature caching, LoSA reaches a $3.2\times$ speedup on HunyuanVideo at a 0.02-point drop, versus 0.32 points for the strongest sparse baseline at comparable speed. Across three video diffusion transformers and speedups up to $3.2\times$, LoSA consistently achieves the best training-free speed-quality trade-off.

↑