AI 中文总结
Vorch-Omni是基于流匹配扩散Transformer的统一多任务视听合成框架,支持10余项视听相关任务,为通用视听生成与操作提供可扩展基础。
AI 中文摘要
生成式视频建模的最新进展已实现多样化生成、基于参考的合成、扩展与编辑,但现有方法常依赖碎片化的特定任务模型。通用模型必须区分异构目标、源与参考信号,以确定生成、保留或用作引导的内容,同时减少任务间的干扰;联合视听生成因引入跨模态的多样化条件与输出配置,进一步加剧了这一挑战。我们提出Vorch-Omni,一种基于任意条件到任意输出形式的视听合成统一多任务框架,可灵活将视频与音频信号视为条件输入或生成目标。令牌级条件掩码与任务标识符区分目标、源内容与参考,位置类型将时间上下文与独立条件分离;为捕捉语义与结构信息,Vorch-Omni采用互补的视觉条件通路:视觉语言模型结合文本指令解释采样帧,视频VAE将条件编码为潜在令牌以直接引导生成。我们还构建了分布式数据管道,用于整理多样化的时间对齐视听片段、生成结构化字幕与元数据,并平衡异构任务分布。Vorch-Omni基于单个流匹配扩散Transformer构建,无需特定任务的架构变更,支持文本到视频、文本到视听、图像及参考条件生成、时间扩展、音频驱动生成、视频转换、视听编辑等10余项任务,为通用视听生成与操作提供了可扩展的基础。
英文摘要
Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks. Joint audio-visual generation further increases this challenge by introducing diverse conditioning and output configurations across modalities. We present Vorch-Omni, a unified multi-task framework for audio-visual synthesis based on an arbitrary-condition-to-arbitrary-output formulation. It flexibly treats video and audio signals as either conditioning inputs or generation targets. Token-level conditioning masks and task identifiers distinguish targets, source content, and references, while position types separate temporal context from independent conditions. To capture semantic and structural information, Vorch-Omni employs complementary visual conditioning pathways: a vision-language model interprets sampled frames with text instructions, and a video VAE encodes conditions into latent tokens for direct guidance. We further build a distributed data pipeline to curate diverse temporally aligned audio-visual clips, generate structured captions and metadata, and balance heterogeneous task distributions. Built on a single flow-matching diffusion transformer without task-specific architectural changes, Vorch-Omni supports over 10 tasks, including text-to-video, text-to-audio-video, image- and reference-conditioned generation, temporal extension, audio-driven generation, video transformation, and audio-visual editing. This unified framework provides a scalable foundation for general-purpose audio-visual generation and manipulation.
CommentsProject Page: https://vorch-project.github.io/Vorch-Omni-project/