发表机构
Baidu; Beijing Institute of Technology(百度; 北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Encore通过自适应信号路由,结合跨块上下文传播与移位位置嵌入参考信号,实现长时长同步视听生成,在VerseBench上性能显著优于现有方法,还支持无限长度跨模态合成。
AI 中文摘要
现有的视听生成方法能生成同步良好的片段,但时长有限;而长视频生成方法虽能通过基于块的迭代合成延长时长,却完全缺乏音频。联合生成长音频视频比单独完成任一任务难度大得多:每个块必须同时保持视频时间连贯性、音频时间连贯性以及跨模态同步,其条件信号通过不同途径进入模型。本研究提出Encore用于长时长同步视听生成,核心思路是将该挑战分解为:(1)通过显式跨块上下文传播的迭代生成处理局部连续性;(2)通过带移位位置嵌入的参考视听信号强制执行全局一致性。基于此设计,提出自适应信号路由(Adaptive Signal Routing, ASR),其在自注意力中引入可学习注意力偏置,在交叉注意力输出上引入可学习残差尺度,使模型能自适应调节各条件信号的影响。Encore经端到端训练用于联合视听生成,推理时通过在整个去噪过程中以真实模态为条件,还支持无限长度的音频转视频和视频转音频合成。在扩展的长视听评估基准VerseBench上的实验表明,Encore在生成质量和时间连贯性上均显著优于现有方法。本论文的代码和数据可在该https网址获取。
英文摘要
Existing audio-video generation methods produce well-synchronized clips but are limited to short durations, while long-video generation methods extend duration through chunk-based iterative synthesis yet lack audio entirely. Generating long audio-video jointly is fundamentally harder than either task alone: each chunk must simultaneously maintain video temporal coherence, audio temporal coherence, and cross-modal synchronization, whose conditioning signals enter the model through different pathways. In this work, we present Encore for long-form synchronized audio-video generation. Our key insight is to factor this challenge into: (1) local continuity handled by iterative generation with explicit cross-chunk context propagation, and (2) global consistency enforced via reference audio-video signals with shifted position embedding. Building on this design, we propose Adaptive Signal Routing (ASR), which introduces learnable attention biases within self-attention and learnable residual scales on cross-attention outputs, enabling the model to adaptively modulate the influence of each conditioning signal. Trained end-to-end for joint audio-video generation, Encore also supports infinite-length audio-to-video and video-to-audio synthesis at inference by conditioning on the ground-truth modality throughout the denoising process. Experiments on our extended VerseBench for long audio-video evaluation demonstrate that Encore significantly outperforms existing methods in both generation quality and temporal coherence. Code and data for this paper are at https://github.com/shaohua-pan/Encore.
CommentsAccepted by SIGGRAPH ASIA 2026