arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21553cs.CV

SANA-Video 2.0:具有注意力残差的混合线性注意力用于高效视频生成

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

Junsong Chen, Jincheng Yu, Yitong Li, Shuchen Xue, Haozhe Liu, Jingyu Xin, Yuyang Zhao, Tian Ye, Zhangjie Wu, Zian Wang, Daquan Zhou, Ping Luo, Song Han, Enze Xie

首次发表
浏览论文内容

中文总结 AI 辅助

SANA-Video 2.0是一种混合视频扩散Transformer,采用混合线性-softmax注意力和块注意力残差,在单个GPU上生成高质量视频,成本大幅降低,实现可扩展长时、高分辨率视频生成,性能与大模型竞争且速度更快。

中文摘要 AI 辅助

我们介绍了SANA-Video 2.0,这是一种在统一架构下实例化的5B和14B规模的混合视频扩散Transformer。旨在在单个GPU上生成高达720p的高质量视频,SANA-Video 2.0在质量上与全softmax视频DiTs相匹配,同时保留线性注意力良好的长序列缩放特性。混合线性-softmax注意力以3:1的比例将门控线性注意力与周期性门控-softmax锚点相结合,避免了二次注意力,恢复了纯线性注意力所缺乏的满秩令牌交互。通过块注意力残差将完整的块摘要路由到后续线性层,实现锚点特征重用并将深层有效秩提高约12%。通过从头开始训练,SANA-Video 2.0直接学习完整的混合模型而非线性化预训练模型,通过降低分辨率的代理研究确定25%的softmax为最佳质量-效率权衡。在40步采样下,SANA-Video 2.0在单个H100上以480p在13.2秒内实现VBench分数84.30,在延迟仅为一小部分的情况下与大得多的softmax视频DiTs竞争。其编译的DiT前向传递在720p/60s时比匹配的全softmax基线快3.2倍,随着视频持续时间差距扩大。全栈Sol-Engine优化进一步加速了这个硬件友好的主干,使5B管道在720p/5s时达到13.06秒,比Wan 2.2-A14B在一个H100上快120倍。总体而言,我们的混合设计以大幅降低的成本恢复了softmax级别的表现力,实现了可扩展的长时、高分辨率视频生成。

英文摘要

We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence scaling of linear attention. To avoid quadratic attention throughout, Hybrid Linear-Softmax Attention combines gated linear attention for O(N)-dominated mixing with periodic gated-softmax anchors at a 3:1 ratio, restoring the full-rank token interactions that pure linear attention lacks. To propagate these refreshed representations across depth, Block Attention Residuals (AttnRes) route completed block summaries into later linear layers, enabling anchor-feature reuse and boosting deep-layer effective rank by ~12%. Through from-scratch training, SANA-Video 2.0 learns the complete hybrid directly rather than linearizing pretrained models, with reduced-resolution proxy studies establishing 25% softmax as the optimal quality-efficiency trade-off. With 40-step sampling, SANA-Video 2.0 achieves a VBench score of 84.30 in 13.2s at 480p on a single H100, remaining competitive with far larger softmax video DiTs at a fraction of the latency. Its compiled DiT forward pass is 3.2x faster than a matched full-softmax baseline at 720p/60s, a gap that expands with video duration. Furthermore, full-stack Sol-Engine optimization (kernel fusion, caching, and sparse attention) accelerates this hardware-friendly backbone by a further 3.58x, bringing the 5B pipeline to 13.06s at 720p/5s and making it 120x faster than Wan 2.2-A14B on one H100. Overall, our hybrid design recovers softmax-level expressiveness at substantially reduced cost, unlocking scalable long, high resolution video generation.

发表机构

  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑