发表机构
University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; Ant Group(中国科学院大学; 中国科学院自动化研究所; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多身份定制文本到视频生成问题,提出GroupVideo框架,融合多模态身份对齐及相关模块,构建高质量数据集,在生成多角色视频时能保持身份一致且动作自然,优于现有方法。
AI 中文摘要
当前身份定制视频生成方法主要限于单身份场景,多身份设置中缺乏明确身份分离机制常导致身份混淆。现有多身份方法通过拼接面部图像扩展单身份框架,常致面部表情和动作不自然。为此引入GroupVideo,基于视频扩散Transformer,融合多模态身份对齐,包括视觉对齐和语义对齐,还引入带空间引导的ID定位模块等。针对多身份视频数据集短缺,构建了20000个视频的数据集。实验表明GroupVideo在生成具有一致身份和自然动作的多角色视频方面优于现有方法。
英文摘要
Current identity customized video generation methodologies are predominantly limited to single-identity scenarios, as the lack of explicit identity separation mechanisms often leads to identity confusion in multi-identity settings. Existing multi-identity approaches, which directly extend single-identity frameworks by concatenating face images as input conditions, frequently result in unnatural facial expressions and motions, manifesting as the "copy-paste" phenomenon. To overcome these limitations, we introduce GroupVideo, a novel framework that leverages multiple individual photographs to generate identitycustomized video. Built upon Video Diffusion Transformers, GroupVideo incorporates multimodal identity alignment: visual alignment jointly encodes multiple face images to provide robust identity references, while semantic alignment introduces a semantic perceiver to enhance the naturalness of motions. An ID localization module with spatial guidance is introduced to address identity blending and enhance identity fidelity, along with bounding box constraints and mask regularization loss, to focus on facial regions and improve training efficiency. In response to the shortage of multi-ID video datasets, we have curated a comprehensive high-quality dataset of 20,000 videos, thereby establishing a crucial resource to advance future research in multi-ID video generation. Extensive experiments demonstrate that GroupVideo outperforms existing methods in generating multi-character videos with consistent identities and natural motions.