arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Vorch-IR:长视频统一多模态身份替换生成模型

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation

Yaole Wang, Xiaoyu Chen, Xin Ma, Yang Ding, Gang Yue, Jingjing Chen, Lin Ma, Yaohui Wang

arXiv 2608.05648首次发表:更新:

发表机构

Vorch Team; Fudan University(Vorch团队; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Vorch-IR是基于LTX2的统一多模态视频身份替换框架,支持单人、双人身份及可选背景替换,通过自动数据构建流水线与时间重叠推理策略,实现长视频生成,在身份保留、动作保真等方面表现出色。

AI 中文摘要

视频身份替换旨在迁移一个或多个主体的身份,同时保留驱动视频的动作、表情和时间结构。现有方法大多针对单人场景,且常需要任务特定的结构控制,如掩码或姿态表示,限制了其在通用多模态编辑系统中的灵活性。多人替换的进展还因配对训练数据稀缺而受阻。我们提出Vorch-IR,这是一个统一框架,支持单人及双人身份替换,可选背景替换,集成在单个模型中。该模型基于LTX2构建,以驱动视频、索引参考图像和文本编辑指令为联合条件。参考图像无需与驱动视频的姿态、布局或空间配置匹配:其作为主体或背景参考的角色通过指令指定。密集视觉条件通过自注意力融合,而视觉语言上下文通过交叉注意力建立语义对应。我们还开发了自动数据构建流水线,为所有四种编辑场景合成配对监督。使用自动指标和成对人工评估的实验表明,在各种场景中,该模型具有强大的身份保留、动作保真度和时间连贯性。时间重叠推理策略还将短片段模型扩展到分钟级生成,无需自回归延续。

英文摘要

Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of a driving video. Existing methods largely target single-person settings and often require task-specific structural controls, such as masks or pose representations, limiting their flexibility in general multimodal editing systems. Progress on multi-person replacement is further constrained by the scarcity of paired training data. We present Vorch-IR, a unified framework that supports single- and dual-person identity replacement, with optional background replacement, in a single model. Built on LTX2, Vorch-IR jointly conditions on a driving video, indexed reference images, and a textual editing instruction. The reference images need not match the pose, layout, or spatial configuration of the driving video: their roles as subject or background references are specified through the instruction. Dense visual conditions are fused through self-attention, while a vision-language context establishes semantic correspondence through cross-attention. We further develop an automatic data construction pipeline that synthesizes paired supervision for all four editing settings. Experiments using automatic metrics and pairwise human evaluation demonstrate strong identity preservation, motion fidelity, and temporal coherence across diverse scenarios. A temporal overlapping inference strategy additionally extends the short-clip model to minute-long generation without autoregressive continuation.

CommentsProject page: https://vorch-project.github.io/Vorch-IR-project/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑