发表机构
Alibaba Group; Alibaba Cloud Computing; Shanghai Jiao Tong University(阿里巴巴集团; 阿里巴巴云计算; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出同步感知加速框架,通过受保护的稀疏注意力策略在保留视听生成质量与同步性的同时提升推理效率。
AI 中文摘要
近期的视听生成模型可在统一扩散过程中合成同步的视频与声音,但其推理成本仍较高,因为长视频 token 序列需要在去噪步骤中重复进行注意力计算。视频生成模型已开发出多种加速技术,包括低位量化、注意力稀疏化和特征蒸馏,但由于这些方法最初是为视频生成设计的,直接应用于视听模型会忽略音频与视频分支之间的交互,因此可能破坏视听同步性。本文提出一种用于高效视听生成的同步感知加速框架。核心观察结果是,双向视听交叉注意力揭示了两个分支之间的结构化交互,高响应通常集中在少数与声音相关的视觉和时间 token 上。通过利用这种交互模式,我们引入了受保护的稀疏注意力策略,该策略在为同步关键 token 保留高保真计算的同时,对冗余注意力进行稀疏化处理。通过在加速过程中明确考虑跨模态依赖关系,我们的方法在保持视频质量、音频质量和视听同步性的同时,提高了推理效率。
英文摘要
Recent audio-visual generation models can synthesize synchronized video and sound in a unified diffusion process, but their inference cost remains high because long video token sequences require repeated attention computation across denoising steps. A variety of acceleration techniques have been developed for video generation models, including low-bit quantization, attention sparsification, and feature caching. However, since these methods are originally designed for video generation, directly applying them to audio-visual models overlooks the interactions between the audio and video branches and may therefore disrupt audio-video synchronization. We present a synchronization-aware acceleration framework for efficient audio-visual generation. Our key observation is that bidirectional audio-video cross-attention reveals structured interactions between the two branches, with high responses often concentrated on a few sound-related visual and temporal regions. Guided by this interaction pattern, we introduce a protected sparse attention strategy that preserves high-fidelity computation for synchronization-critical tokens while sparsifying redundant attention interactions. By explicitly accounting for cross-modal dependence during acceleration, our method improves inference efficiency while keeping video quality, audio quality, and audio-video synchronization.