arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05776cs.CV

Vorch-Director:基于噪声感知误差校正的交互式世界故事模型

Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification

Lisai Zhang, Yidi Wu, Qi Liu, Xin Ma, Yang Ding, Gang Yue, Siqian Yang, Jingyuan Chen, Lin Ma, Yaohui Wang

首次发表
浏览论文内容

中文总结 AI 辅助

Vorch-Director是一种噪声感知残差校正策略,基于LTX-2扩散Transformer,可提升视听长视频生成的稳定性与保真度,支持多镜头、多主体的参考引导生成。

中文摘要 AI 辅助

自回归续成为分钟级视听生成提供了自然路径,其通过反复扩展以先前生成的视频和音频为条件的短窗口生成器实现。然而,模型在干净的真实历史上训练,而推理依赖于自身生成的历史,累积误差会导致身份漂移、过度平滑以及视听不同步。近期方法通过重用预测残差作为合成损坏来减少这种不匹配,但我们观察到,残差校正的有效性关键取决于产生残差的流匹配噪声水平。我们提出Vorch-Director,一种噪声水平感知的残差校正策略,该策略将每个残差与其产生的噪声水平关联,并在训练期间注入匹配噪声 regime 的残差。通过将注入的误差与去噪过程对齐,Vorch-Director生成更真实的自回归历史,同时保留高效的教师强制训练。基于视听LTX-2扩散Transformer,Vorch-Director进一步引入任务嵌入以区分历史视频、参考图像和目标视频,为长程生成实现统一条件。结合干净的条件汇和混合任务训练,Vorch-Director支持多镜头、多主体、参考引导的视听长视频生成。我们在ST-Bench上评估Vorch-Director,并引入新的长程视听基准,包含质量漂移和长程一致性指标。大量实验表明,与强基线相比,Vorch-Director在稳定性和视听保真度上均有提升。

英文摘要

Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditioned on previously generated video and audio. However, models are trained on clean ground-truth histories, while inference relies on their own generated histories, where accumulated errors cause identity drift, over-smoothing, and audio-visual desynchronization. Recent methods reduce this mismatch by reusing prediction residuals as synthetic corruption, but we observe that the effectiveness of residual correction critically depends on the flow-matching noise level at which residuals are produced. We propose Vorch-Director, a noise-level-aware residual correction strategy that associates each residual with its originating noise level and injects residuals from matched noise regimes during training. By aligning injected errors with the denoising process, Vorch-Director produces more realistic autoregressive histories while retaining efficient teacher-forcing training. Built on the audio-visual LTX-2 diffusion transformer, Vorch-Director further introduces task embeddings to distinguish historical video, reference images, and target video, enabling unified conditioning for long-horizon generation. Together with a clean conditioning sink and mixed-task training, Vorch-Director supports multi-shot, multi-subject, reference-guided audio-visual long-video generation. We evaluate Vorch-Director on ST-Bench and introduce a new long-horizon audio-visual benchmark with metrics for quality drift and long-range consistency. Extensive experiments demonstrate improved stability and audio-visual fidelity over strong baselines.

发表机构

  • Vorch Team(Vorch团队)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑