Prism:用于原生2K联合视频-音频生成模型训练的动态稀疏注意力
Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training
浏览论文内容
中文总结 AI 辅助
Prism提出动态稀疏注意力框架,通过时空宏区域和自适应块形状,实现2K联合视频-音频生成模型的原生训练,获得2.5倍加速并提升生成质量。
中文摘要 AI 辅助
以更高分辨率原生训练联合视频-音频生成模型,使其能够学习更丰富的视觉细节和更锐利的运动动态。然而,全注意力机制会产生二次方计算成本,并且随着分辨率的提高,注意力会分散到日益冗余的令牌上,从而稀释了信息丰富内容的学习信号,并扰乱了预训练先验。现有的稀疏注意力方法要么针对免训练加速,要么忽视了联合视频-音频数据的独特结构,其中跨模态交互本质上集中在产生声音的区域周围。为了解决这个问题,我们提出了Prism,一个用于在2K分辨率下原生训练联合视频-音频生成模型的动态稀疏注意力框架。具体而言,Prism将令牌序列组织成时空宏区域,使注意力结构能够适应局部内容。对于每个区域,它通过沿通道的视频特征方差和来自音频到视频交叉注意力的特征范数来估计局部信息结构,共同捕捉视觉内容如何方向性变化以及音频对每个视觉区域的影响强度。基于这些信号,Prism动态地为每个区域分配定制的块形状,沿视觉内容快速变化和强音视频耦合的轴应用更精细的划分。这促进了每个块内令牌保持语义连贯性,使得块级特征能够同时捕捉视觉内容和联合视频-音频交互模式。Prism进一步采用混合块选择策略来动态确定每个查询的稀疏性。实验表明,与全注意力相比,Prism实现了2.5倍的训练加速,同时在生成质量上超越了全注意力。
英文摘要
Natively training joint video-audio generation models at higher resolutions empowers them to learn richer visual details and sharper motion dynamics. However, full attention incurs quadratic cost and, as resolution increases, spreads attention over increasingly redundant tokens, diluting learning signals for informative content and disrupting pretrained priors. Existing sparse attention methods either target training-free acceleration or overlook the unique structure of joint video-audio data, where cross-modal interactions are inherently concentrated around sound-producing regions. To address this, we propose Prism, a dynamic sparse attention framework for natively training joint video-audio generation models at 2K. In particular, Prism organizes the token sequence into spatiotemporal macro-zones, enabling the attention structure to adapt to local content. For each zone, it estimates local information structure via video feature variance along the channel and feature norms from the audio-to-video cross-attention, jointly capturing how visual content varies directionally and how strongly audio influences each visual region. Based on these signals, Prism dynamically assigns a tailored block shape to each zone, applying finer partitioning along axes of rapid visual content variation and strong audio-visual coupling. This encourages tokens within each block to remain semantically coherent, allowing block-level features to capture both visual content and joint video-audio interaction patterns. Prism further adopts a hybrid block selection strategy to dynamically determine per-query sparsity. Experiments show that Prism achieves 2.5$\times$ training speedup compared to full attention, while surpassing it in generation quality.
发表机构
- Fudan University(复旦大学)
- Tencent Hunyuan Foundation Model Team(腾讯混元基础模型团队)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。