arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Vera:身份忠实的人类主体到视频生成

Vera: Identity-Faithful Human Subject-to-Video Generation

Yulong Xu, Xinyue Liu, Shujuan Li, huafeng shi, Yan Zhou, Jiwen Liu, Xintao Wang, Yu Shen Liu, Huaibo Huang

arXiv 2607.20247首次发表:更新:

发表机构

Kuaishou Technology; Institute of Automation, Chinese Academy of Sciences; Tsinghua University(快手科技; 中国科学院自动化研究所; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对主体到视频生成中人类身份一致性不足问题,提出Vera框架,通过构建百万对身份对齐数据集,引入身份聚焦掩码监督和参考感知分层注意力两种设计,提升了身份一致性、主体绑定及运动自然性等。

AI 中文摘要

主体到视频(S2V)生成在跨不同类别保留参考主体方面取得了重大进展,但以人类为中心的生成中通用的主体一致性仍然不足。在多人场景中问题更严重,身份角色绑定错误会导致主体混淆等。我们提出了Vera,一个用于单人和多人生成的统一的以人类为中心的S2V框架。首先通过人物级跨剪辑检索构建了一个百万对身份对齐的人类图像 - 视频数据集。在此数据集基础上,Vera引入了两种互补设计。身份聚焦掩码监督(IFMS)加强身份感知学习,参考感知分层注意力(RALA)调节视频令牌与参考身份线索的交互。大量实验表明,Vera提高了人类身份一致性、多人主体绑定和运动自然性,同时减少了身份混淆和过度的参考图像复制。

英文摘要

Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across diverse categories, yet generic subject consistency remains insufficient for human-centric generation. A video may appear globally consistent while identity-critical human details still drift across frames, poses, and interactions. This issue becomes more severe in multi-person scenarios, where incorrect identity-role binding leads to subject confusion, attribute swapping, and excessive copying of reference-specific appearance cues. We propose Vera, a unified human-centric S2V framework for single- and multi-person generation. We first construct a million-pair identity-aligned human image-video dataset through person-level cross-clip retrieval, providing explicit identity correspondence and diverse references. Built on this dataset, Vera introduces two complementary designs. Identity-Focal Masked Supervision (IFMS) strengthens identity-aware learning with spatially focused supervision while reducing interference from irrelevant artifacts. Reference-Aware Layer-wise Attention (RALA) regulates how video tokens interact with reference identity cues in the DiT backbone, preserving stable identity anchors and enhancing layer-aware identity readout. Extensive experiments demonstrate that Vera improves human identity consistency, multi-person subject binding, and motion naturalness, while reducing identity confusion and excessive reference-image copying.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑