arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2505.22918cs.CV

Re-ttention:基于注意力统计重塑的超稀疏视觉生成

Re-ttention: Ultra Sparse Visual Generation via Attention Statistical Reshape

  • ECE Department University of Alberta(阿尔伯塔大学电子工程系)
  • Division of CSE Louisiana State University(路易斯安那州立大学计算机科学与工程学院)
  • Huawei Technologies Edmonton, Alberta, Canada(华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Ruichen Chen, Keith G. Mills, Liyao Jiang, Chao Gao, Di Niu

更新

AI总结:

针对扩散Transformer注意力机制复杂度高、现有稀疏注意力在极高稀疏度下画质损失大的问题,提出Re-ttention方法,通过注意力统计重塑实现超稀疏视觉生成,推理仅需3.1%的token且性能优于同类方法。

AI中文摘要:

扩散Transformer(DiT)已成为生成视频、图像等高质量视觉内容的事实标准模型。其一大瓶颈是注意力机制,复杂度随分辨率和视频长度呈二次方增长。减轻这一负担的一种合理思路是稀疏注意力,即计算时仅包含部分token或patch。然而,现有技术在极高稀疏度下无法保持视觉质量,甚至可能带来不可忽视的计算开销。为解决这一问题,我们提出Re-ttention,该方法利用扩散模型的时间冗余性克服注意力机制内的概率归一化偏移,为视觉生成模型实现极高稀疏度的注意力。具体而言,Re-ttention基于先前的softmax分布历史重塑注意力分数,从而在极高稀疏度下保持完整二次注意力的视觉质量。在CogVideoX、PixArt DiT等T2V/T2I模型上的实验结果表明,Re-ttention在推理时仅需低至3.1%的token,性能优于FastDiTAttn、Sparse VideoGen、MInference等现有方法。

英文摘要:

Diffusion Transformers (DiT) have become the de-facto model for generating high-quality visual content like videos and images. A huge bottleneck is the attention mechanism where complexity scales quadratically with resolution and video length. One logical way to lessen this burden is sparse attention, where only a subset of tokens or patches are included in the calculation. However, existing techniques fail to preserve visual quality at extremely high sparsity levels and might even incur non-negligible compute overheads. To address this concern, we propose Re-ttention, which implements very high sparse attention for visual generation models by leveraging the temporal redundancy of Diffusion Models to overcome the probabilistic normalization shift within the attention mechanism. Specifically, Re-ttention reshapes attention scores based on the prior softmax distribution history in order to preserve the visual quality of the full quadratic attention at very high sparsity levels. Experimental results on T2V/T2I models such as CogVideoX and the PixArt DiTs demonstrate that Re-ttention requires as few as 3.1% of the tokens during inference, outperforming contemporary methods like FastDiTAttn, Sparse VideoGen and MInference.

补充信息

↑