通过并行适配器组合实现实时音视频联合生成
Real-Time Joint Audio-Video Generation by Parallel Adapter Composition
浏览论文内容
中文总结 AI 辅助
针对实时交互式音视频生成,提出在冻结骨干上并行训练因果与少步适配器,利用近似正交更新直接相加组合,实现无需联合训练的流式生成,质量媲美教师模型。
中文摘要 AI 辅助
部署用于实时、交互式生成的联合音视频扩散Transformer通常需要两项关键修改:块自回归注意力,使得帧可以在整个片段完成之前输出;以及少步采样,使每个块的生成成本降低。传统上,流式视频文献通过链式流水线获得这两种能力:首先将双向教师模型蒸馏为因果学生模型,再蒸馏为少步模型,或按相反顺序进行。这种链的每个阶段都会微调前一阶段产生的权重,因此后续目标可能破坏先前的能力。借鉴模型合并的思想,我们证明在打包的音视频骨干网络上,这两种能力可以并行获取。一个因果适配器针对冻结的骨干网络进行训练,而一个现成的少步适配器提供少步能力。由于两者作用于不同的功能轴,我们预测并验证了它们的权重更新方向近似正交,无需在训练期间施加显式正交约束。正交更新应能无干扰地组合,因此并行组合是一种直和。两个适配器在推理时简单相加,无需联合训练,即可产生少步、流式的音视频,其图像质量与双向教师模型相当。与链式基线相比,组合模型在大多数指标上达到或超越它们,使并行组合成为一种实用方法。所得到的流式系统能够实时生成联合音视频,在480×832分辨率下无需量化即可达到约26帧/秒,并可持续生成30秒且图像质量稳定。
英文摘要
Deploying a joint audio-video diffusion transformer for real-time, interactive generation normally requires two essential modifications: block-autoregressive attention, so frames can be emitted before the whole clip is finished, and few-step sampling, so each block is cheap. Conventionally, the streaming video literature obtains both capabilities from a chained pipeline. It first distills a bidirectional teacher into a causal student, then into a few-step one, or proceeds in reverse order. Each stage of such a chain fine-tunes the weights the previous one produced, so a later objective can undo an earlier capability. Following the idea of model merging, we show that on a packed audio-video backbone the two capabilities can be acquired in parallel. A causal adapter is trained against the frozen backbone, and an off-the-shelf few-step adapter provides the few-step capability. As the two edit different functional axes, we predict, and then verify, that their weight-update directions are near-orthogonal, without any explicit orthogonality constraint during training. Orthogonal updates should combine without interfering, so parallel composition is a direct sum. The two adapters are simply added at inference, with no joint training, yielding few-step, streaming audio-video whose image quality tracks the bidirectional teacher. Compared to the chained baselines, the composed model matches or beats them on most metrics, making parallel composition a practical approach. The resulting streaming system generates joint audio-video in real time, $\approx$26 fps at $480\times832$ without quantization, and sustains 30 s of continuous generation with stable image quality. Demo Page: https://pac-demo-2027.github.io/demo/
发表机构
- LIGHTSPEED
机构由 AI 辅助整理,请以论文原文为准。