arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Video DeltaNet:用于直播视频生成的视频原生混合注意力

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Haocheng Xi, Yiming Xie, Hexu Zhao, Yiwen Zhang, Michael Liu, Thomas Creavin, Kurt Keutzer, Xiuyu Li, Zhaoyang Lv, Chenfeng Xu, Haiwen Feng

arXiv 2609.20744首次发表:更新:

发表机构

Impossible, Inc.(Impossible 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Video DeltaNet,一种结合局部Softmax注意力与双向线性记忆的视频原生混合注意力机制,用于直播视频生成,在保持高质量的同时显著加速扩散模型去噪过程。

AI 中文摘要

视频扩散模型在去噪过程中反复处理长时间跨度的时空标记序列,这使得注意力机制成为主要的计算瓶颈。线性注意力提供了一种有吸引力的替代方案,并已在近期的大型语言模型中被广泛采用,但将其直接应用于视频模型往往无法保留高质量生成所需的细粒度交互。我们提出了Video DeltaNet(VDN),它将局部Softmax注意力与双向线性记忆相结合,以处理长距离视频上下文。其线性分支引入了Video Delta Attention(VDA),该机制通过每帧整合其空间标记来更新一次记忆。独立的输出投影和可学习的门控对两个分支进行校准,同时分阶段的教师对齐策略逐步将新通路引入预训练模型。我们在MiniMax H3上实例化VDN,将混合注意力应用于视频到视频的交互,同时保留Softmax用于涉及文本或音频的交互。通过八步蒸馏和优化的SGLang服务栈,VDN-H3在八块NVIDIA B200 GPU上完成14.3秒、768p视频的DiT去噪仅需6.70秒,相比同GPU数量下50步的密集H3基线实现了14.5倍的加速。

英文摘要

Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models often fails to preserve the fine-grained interactions required for high-quality generation. We present Video DeltaNet (VDN), which combines local Softmax attention with bidirectional linear memory for long-range video context. Its linear branch introduces Video Delta Attention (VDA), which updates memory once per frame by jointly incorporating its spatial tokens. Separate output projections and learnable gates calibrate the two branches, while a staged teacher-alignment recipe progressively introduces the new pathway into pretrained models. We instantiate VDN on MiniMax H3, applying the hybrid to video-to-video interactions while retaining Softmax for interactions involving text or audio. With eight-step distillation and an optimized SGLang serving stack, VDN-H3 completes DiT denoising for a 14.3-second, 768p video in 6.70 seconds on eight NVIDIA B200 GPUs, corresponding to a 14.5x speedup over the 50-step dense H3 baseline on the same GPU count. GitHub code available at: https://github.com/OpenVDN/vdn-minimax-h3. Weights available at: https://huggingface.co/OpenVDN/vdn-minimax-h3

CommentsGitHub code available at: https://github.com/OpenVDN/vdn-minimax-h3. Weights available at: https://huggingface.co/OpenVDN/vdn-minimax-h3

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑