arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05663cs.CVcs.SD

Vorch-Streamer:将人类音视频生扩展至实时长格式流式生成

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming

Menglin Han, Yang Ding, Yulei Lu, Haoran Yu, Xin Ma, Junyi Chen, Zhangkai Ni, Lin Ma, Yaohui Wang

首次发表
浏览论文内容

中文总结 AI 辅助

Vorch-Streamer是一款后训练框架,通过混合训练、自强制与DMD蒸馏等技术,实现了超24 FPS的实时长格式T2AV流式音视频生成,兼具音唇同步性与身份保留能力。

中文摘要 AI 辅助

实时长格式虚拟形象音视频生成需要实现因果连续合成,同时保持音视频同步性与视觉一致性。将预训练的双向模型适配至该场景会面临两大关键难题:其一,自回归地复用生成块作为上下文会产生暴露偏差,导致错误与视觉漂移在长期生成过程中不断累积;其二,当仅能获取有限的局部音视频上下文时,全局语音表述无法指示因果生成器应生成的下一语音部分。我们提出Vorch-Streamer,这一后训练框架可应对上述挑战,实现实时长格式文本到音视频(T2AV)的流式生成。我们构建了包含8万段、时长12-21秒的虚拟形象片段的合成语料库,先通过混合教师强制与扩散强制训练因果生成器;再结合长 horizon 自强制与DMD蒸馏,让模型接触自身的生成分布,同时保留预训练双向教师模型的质量。为显式控制语音进度,外部语言模型会预测离散的25Hz语音规划标记,其连续特征用于条件化音频扩散分支,并使每个因果块与应生成的内容对齐。在有限因果上下文与四步去噪的条件下,Vorch-Streamer以27.12 FPS的速率从文本联合生成音频与视频,超过24 FPS的实时播放速率,同时在长格式生成中保持了有竞争力的音唇同步性与较强的身份保留能力。

英文摘要

Real-time long-form avatar audio-video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Adapting a pretrained bidirectional model to this setting presents two key dilemmas. First, autoregressively reusing generated blocks as context creates exposure bias, causing errors and visual drift to accumulate over long rollouts. Second, a global speech utterance does not indicates a causal generator which portion should be spoken next when only limited local audio-video context is available. We present Vorch-Streamer, a post-training framework that addresses these challenges and enables real-time long-form Text-to-Audio-Video (T2AV) streaming. We construct a synthetic corpus of 80K avatar clips spanning 12-21 seconds and first train a causal generator with mixed Teacher Forcing and Diffusion Forcing. We then apply long-horizon Self Forcing with DMD distillation, exposing the model to its own rollout distribution while preserving the quality of the pretrained bidirectional teacher. To explicitly control speech progression, an external language model predicts discrete 25-Hz speech-planning tokens, whose continuous features condition the audio diffusion branch and align each causal block with the content it should speak. With bounded causal context and four-step denoising, Vorch-Streamer jointly generates audio and video from text at 27.12 FPS, exceeding the 24-FPS real-time playback rate while maintaining competitive audio-lip synchronization and strong identity preservation over long-form generation.

发表机构

  • Vorch Team(Vorch团队)
  • Tongji University(同济大学)
  • Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑