SANA-Video 2.0:具有注意力残差的混合线性注意力用于高效视频生成
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
浏览论文内容
中文总结 AI 辅助
SANA-Video 2.0是一种混合视频扩散Transformer,采用混合线性-softmax注意力和块注意力残差,在单个GPU上生成高质量视频,成本大幅降低,实现可扩展长时、高分辨率视频生成,性能与大模型竞争且速度更快。
中文摘要 AI 辅助
我们介绍了SANA-Video 2.0,这是一种在统一架构下实例化的5B和14B规模的混合视频扩散Transformer。旨在在单个GPU上生成高达720p的高质量视频,SANA-Video 2.0在质量上与全softmax视频DiTs相匹配,同时保留线性注意力良好的长序列缩放特性。混合线性-softmax注意力以3:1的比例将门控线性注意力与周期性门控-softmax锚点相结合,避免了二次注意力,恢复了纯线性注意力所缺乏的满秩令牌交互。通过块注意力残差将完整的块摘要路由到后续线性层,实现锚点特征重用并将深层有效秩提高约12%。通过从头开始训练,SANA-Video 2.0直接学习完整的混合模型而非线性化预训练模型,通过降低分辨率的代理研究确定25%的softmax为最佳质量-效率权衡。在40步采样下,SANA-Video 2.0在单个H100上以480p在13.2秒内实现VBench分数84.30,在延迟仅为一小部分的情况下与大得多的softmax视频DiTs竞争。其编译的DiT前向传递在720p/60s时比匹配的全softmax基线快3.2倍,随着视频持续时间差距扩大。全栈Sol-Engine优化进一步加速了这个硬件友好的主干,使5B管道在720p/5s时达到13.06秒,比Wan 2.2-A14B在一个H100上快120倍。总体而言,我们的混合设计以大幅降低的成本恢复了softmax级别的表现力,实现了可扩展的长时、高分辨率视频生成。
英文摘要
We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence scaling of linear attention. To avoid quadratic attention throughout, Hybrid Linear-Softmax Attention combines gated linear attention for O(N)-dominated mixing with periodic gated-softmax anchors at a 3:1 ratio, restoring the full-rank token interactions that pure linear attention lacks. To propagate these refreshed representations across depth, Block Attention Residuals (AttnRes) route completed block summaries into later linear layers, enabling anchor-feature reuse and boosting deep-layer effective rank by ~12%. Through from-scratch training, SANA-Video 2.0 learns the complete hybrid directly rather than linearizing pretrained models, with reduced-resolution proxy studies establishing 25% softmax as the optimal quality-efficiency trade-off. With 40-step sampling, SANA-Video 2.0 achieves a VBench score of 84.30 in 13.2s at 480p on a single H100, remaining competitive with far larger softmax video DiTs at a fraction of the latency. Its compiled DiT forward pass is 3.2x faster than a matched full-softmax baseline at 720p/60s, a gap that expands with video duration. Furthermore, full-stack Sol-Engine optimization (kernel fusion, caching, and sparse attention) accelerates this hardware-friendly backbone by a further 3.58x, bringing the 5B pipeline to 13.06s at 720p/5s and making it 120x faster than Wan 2.2-A14B on one H100. Overall, our hybrid design recovers softmax-level expressiveness at substantially reduced cost, unlocking scalable long, high resolution video generation.
发表机构
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。