SQuad:用于高效视频生成的次二次注意力蒸馏
SQuad: Sub-Quadratic Attention Distillation for Efficient Video Generation
浏览论文内容
中文总结 AI 辅助
针对视频扩散Transformer自注意力的二次复杂度瓶颈,提出SQuad次二次注意力蒸馏框架,通过两阶段蒸馏适配预训练模型,在保持生成质量的同时大幅提升效率。
中文摘要 AI 辅助
视频扩散Transformer(DiTs)的大部分计算消耗集中在自注意力操作中,该操作的成本随潜在令牌数量n呈二次增长,为O(n²)。在视频生成任务中,令牌数量庞大,因此该操作主导了运行时间和内存消耗,从而限制了可生成内容的分辨率和时长。自注意力的线性O(n)和低秩O(nk)替代方案将完整的softmax QK^T替换为更廉价的内核,但往往无法恢复原始模型的表达能力,存在明显的质量差距。受此启发,我们提出SQuad,这是一个次二次注意力蒸馏框架,其生成的注意力复杂度为O(n√n),自然平衡了效率与表达能力之间的权衡。我们没有从头训练自己的视频DiT(这成本过高),而是通过两个阶段的蒸馏将预训练的完整softmax自注意力DiT适配到我们提出的SQuad-Attention中:流匹配监督微调(SFT),以及改进的分布匹配蒸馏(DMD2),后者还能提升采样效率。在Wan~2.2 5B文本到视频模型上,SQuad在VBench指标上与二次复杂度的教师模型相当(83.20 vs 83.08),同时将每步每块注意力的浮点运算量(FLOPs)削减了约67倍,注意力延迟降低了约11倍,端到端DiT延迟降低了2倍,且仅需6次神经函数评估(NFEs)即可生成视频,而默认设置为100次。
英文摘要
Video Diffusion Transformers (DiTs) spend most of their compute inside the Self-Attention operation, whose cost grows quadratically, $\mathcal{O}(n^2)$, with the number of latent tokens $n$. For the task of video generation, the token count is large, so this term dominates runtime and memory, and thereby caps the resolution and duration we can generate. Linear $\mathcal{O}(n)$ and low-rank $\mathcal{O}(nk)$ surrogates of Self-Attention trade the full softmax $QK^T$ for cheaper kernels, but rarely recover the original's expressivity, leaving a stubborn quality gap. Motivated by this, we propose SQuad, a Sub-Quadratic Attention Distillation framework that achieves a complexity of $\mathcal{O}(n\sqrt{n})$ in the resulting distilled Attention, naturally balancing the efficiency v/s expressivity trade-off. Instead of training our own Video DiT from scratch, which is prohibitively expensive, we fit a pretrained full softmax Self-Attention DiT into our proposed SQuad-Attention one by distilling the former in two stages: Flow-Matching Supervised Fine-Tuning (SFT), followed by improved Distribution Matching Distillation (DMD2) which additionally makes the sampling more efficient. On the Wan~2.2 5B text-to-video model, SQuAD matches the quadratic teacher on VBench ($83.20$ v/s $83.08$) while cutting the per-step per-block attention FLOPs by $\sim$$67\times$ and attention latency by $\sim$$11\times$, and end-to-end DiT latency by 2$\times$, all while also generating a video in only $6$ Neural Functional Evaluations (NFEs) instead of the default $100$.