发表机构
University of Memphis; University of Pennsylvania; University of Iowa(孟菲斯大学; 宾夕法尼亚大学; 爱荷华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文对比研究直接、并行、顺序三种无训练多主体图像到视频生成范式,通过实证评估明确各范式的优劣势,为可控多主体视频生成系统设计提供实用见解。
AI 中文摘要
文本条件图像到视频(I2V)生成已快速发展,但生成包含多个主体的视频仍具挑战性。模型必须同时保留每个主体的外观、分配不同的动作,并维持连贯的空间与时间交互。本文对无训练多主体I2V生成的三种代表性范式开展系统研究:直接式、并行式与顺序式生成。直接式生成将预训练I2V模型应用于完整参考图像与提示,需联合合成所有主体与动作。并行式与顺序式生成则将参考图像和提示分解为主体特定的视觉与文本条件。并行式生成独立合成每个主体,随后组合生成的视频,降低了各生成步骤的复杂度,但代价是主体间上下文较弱。顺序式生成先合成背景视频,再逐步引入单个主体,这保留了累积的场景上下文,但对主体顺序敏感且存在误差传播。我们在多样的多主体场景中对三种范式进行实证评估,对比外观保留度、动作保真度、时间一致性与主体间连贯性,同时刻画其各自的失败模式。研究结果揭示了每种范式的优势与局限,为设计可控的多主体视频生成系统提供实用见解。
英文摘要
Text-conditioned image-to-video (I2V) generation has advanced rapidly, yet generating videos with multiple subjects remains challenging. A model must simultaneously preserve the appearance of each subject, assign distinct motions, and maintain coherent spatial and temporal interactions. This paper presents a systematic study of three representative paradigms for training-free multi-subject I2V generation: direct, parallel, and sequential generation. Direct generation applies a pretrained I2V model to the complete reference image and prompt, requiring all subjects and motions to be synthesized jointly. Parallel and sequential generation instead decompose the reference image and prompt into subject-specific visual and textual conditions. Parallel generation synthesizes each subject independently and subsequently composes the resulting videos, reducing the complexity of each generation step at the cost of weaker inter-subject context. Sequential generation first synthesizes a background video and then progressively introduces individual subjects. This preserves accumulated scene context but introduces sensitivity to subject ordering and error propagation. We empirically evaluate the three paradigms across diverse multi-subject scenes, comparing appearance preservation, motion fidelity, temporal consistency, and inter-subject coherence, while also characterizing their distinct failure modes. Our findings reveal the strengths and limitations of each paradigm and offer practical insights for designing controllable multi-subject video generation systems.
CommentsACM Multimedia Workshop 2026