计算何处重要:用于高效视频扩散的异构注意力
Where Compute Matters: Heterogeneous Attention for Efficient Video Diffusion
浏览论文内容
中文总结 AI 辅助
针对视频扩散中自注意力二次方成本问题,提出异构注意力机制HetA-DiT,按去噪难度动态分配计算,仅约20%标记用密集注意力,在保持生成质量的同时提升效率。
中文摘要 AI 辅助
高效视频生成需要降低长时空标记序列上自注意力的二次方成本。现有的高效注意力方法通常对每个标记应用相同的计算模式,尽管去噪难度在视频区域间差异显著,并在生成过程中不断演变。我们引入了HetA-DiT,一种异构注意力机制,根据标记难度自适应地分配计算。一个轻量级不确定性分支预测逐标记的去噪难度估计,用于将不确定的标记路由到密集全局注意力,同时用高效局部注意力处理更可靠的标记。由此产生的路由是内容和时间步自适应的,在最需要的地方保留全局上下文,并提供一个单一参数来控制质量-效率权衡。HetA-DiT与少步分布匹配蒸馏兼容,并通过重用前一个去噪步骤的不确定性估计,在推理时不引入额外的Transformer评估。我们在DMD蒸馏的Wan2.2-5B和Wan2.1-1.3B模型上评估该方法。在VBench、VBench-2.0和人类偏好评估中,HetA-DiT保持了有竞争力的生成质量,同时仅将约20%的标记路由到密集注意力。
英文摘要
Efficient video generation requires reducing the quadratic cost of self-attention over long spatio-temporal token sequences. Existing efficient-attention methods typically apply the same computation pattern to every token, even though denoising difficulty varies substantially across video regions and evolves throughout the generation process. We introduce HetA-DiT, a heterogeneous attention mechanism that adaptively allocates computation according to token difficulty. A lightweight uncertainty branch predicts a token-wise estimate of denoising difficulty, which is used to route uncertain tokens through dense global attention while processing more reliable tokens with efficient local attention. The resulting routing is content- and timestep-adaptive, retains global context where it matters most, and provides a single parameter for controlling the quality-efficiency trade-off. HetA-DiT is compatible with few-step distribution-matching distillation and introduces no additional Transformer evaluation at inference time by reusing uncertainty estimates from the preceding denoising step. We evaluate the method on DMD-distilled Wan2.2-5B and Wan2.1-1.3B models. Across VBench, VBench-2.0, and human preference evaluation, HetA-DiT maintains competitive generation quality while routing only approximately 20% of tokens through dense attention.