发表机构
Vorch Team; Harbin Institute of Technology, Shenzhen; Tongji University(Vorch团队; 哈尔滨工业大学(深圳); 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Vorch-Human提出统一的多任务以人为中心生成框架,通过双流扩散Transformer和两级数据流水线,实现短时及五分钟长视频的强身份保持、音视频同步与时间稳定性。
AI 中文摘要
以人为中心的音视频生成涵盖多个紧密相关的任务:根据驱动语音为人物制作动画、根据声音参考联合生成语音和视频、以及根据配对的外观和声音参考合成场景。现有系统通常使用单独的模型来解决这些任务,尽管它们共享相同的目标模态,并且主要区别在于提供哪些观测作为条件。我们提出了Vorch-Human,一个基于双流音视频扩散Transformer的统一以人为中心的生成框架。Vorch-Human在传统的噪声音频/噪声视频接口上增加了干净的条件音频和条件视频令牌组。每个令牌的任务嵌入、时间位置类型、条件掩码以及共享的多模态提示编码器使得驱动语音、音色示例、首帧和主体图像能够在单个模型中表达。为了提供该接口所需的监督,我们开发了一个两级数据流水线。第一级使用语音识别、人声分离、人脸检测与跟踪、活动说话人和同步模型、音频/视觉说话人聚类以及多模态字幕校正来分析每个片段;它生成主体索引的语音、外观和音色注释。第二级将来自同一源视频的不同片段中的同一个人关联起来,并在人脸、身体、质量、姿态和视觉语言验证后挖掘身份和着装一致的参考图像。最后,我们通过使用干净潜在前缀进行训练,并在推理时使用相同的冻结前缀循环,将Vorch-Human适应于长格式音频驱动的生成。每个片段仅贡献其新生成的后缀,减少了边界不连续性和长时程身份漂移。在短片段和五分钟生成上的实验展示了强大的身份保持、音视频同步和时间稳定性。
英文摘要
Human-centric audio-visual generation spans several closely related tasks: animating a person from driving speech, jointly generating speech and video from a voice reference, and synthesizing a scene from paired appearance and voice references. Existing systems commonly solve these tasks with separate models, even though they share the same target modalities and differ mainly in which observations are provided as conditions. We present Vorch-Human, a unified human-centric generation framework built on a dual-stream audio-video diffusion transformer. Vorch-Human augments the conventional noisy audio/noisy video interface with clean condition-audio and condition-video token groups. Per-token task embeddings, temporal position types, condition masks, and a shared multimodal prompt encoder allow driving speech, timbre examples, first frames, and subject images to be expressed within one model. To supply the supervision required by this interface, we develop a two-level data pipeline. Level 1 analyzes each clip with speech recognition, vocal separation, face detection and tracking, active-speaker and synchronization models, audio/visual speaker clustering, and multimodal caption correction; it produces subject-indexed speech, appearance, and timbre annotations. Level 2 links the same person across clips from a common source video and mines identity- and outfit-consistent reference images after face, body, quality, pose, and vision-language verification. Finally, we adapt Vorch-Human to long-form audio-driven generation by training with clean latent prefixes and using the same frozen-prefix recurrence at inference. Each segment contributes only its newly generated suffix, reducing boundary discontinuity and long-horizon identity drift. Experiments on short and five-minute generation demonstrate strong identity preservation, audio-visual synchronization, and temporal stability.
CommentsProject page: https://vorch-project.github.io/Vorch-Human-Project/