WithEveryone:面向群体图像生成的统一规划与身份定位
WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
浏览论文内容
中文总结 AI 辅助
针对多人物群体图像生成的身份保留难题,提出WithEveryone框架,通过布局定位身份损失等技术,提升了人脸相似度并降低伪影,实现高身份覆盖率与低重复率。
中文摘要 AI 辅助
当场景需包含多名指定人物时,身份保留式图像生成的可靠性会显著下降。除了要保留每个身份外,模型还需将每个参考身份与不同的人物、位置绑定,而训练阶段的身份损失函数必须在多个含噪声的预测人脸间建立对应关系。我们提出WithEveryone,这是一个可生成包含最多10个参考身份的群体图像的统一框架。WithEveryone将每个选定身份作为带地址的token注入,预测结构化的身份-布局规划,并将该规划渲染为视觉条件。其核心目标是布局定位身份损失(Layout-Grounded ID Loss),该损失函数使用带注释的人脸区域直接监督目标身份,避免基于嵌入的人脸匹配的不稳定性;身份表示强制(ID Representation Forcing)则在图像合成前为每个身份训练预测模型。在一个身份不重叠的基准测试中,WithEveryone实现了最高的目标-上下文身份相似度,将人脸相似度从GPT-Image-2的0.462提升至0.499,同时将复制粘贴伪影从0.169降至0.055。它还覆盖了97.3%的请求身份,重复率仅为2.8%。这些结果表明,显式的身份-布局定位可使身份保留式生成扩展至更大的群体,无需依赖直接的参考人脸复制。
英文摘要
Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as an addressed token, predicts a structured identity--layout plan, and renders the plan as a visual condition. Its key objective, Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching; ID Representation Forcing additionally trains a prediction for each identity before image synthesis. On an identity-disjoint benchmark, WithEveryone achieves the highest target-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3\% of the requested identities with a duplicate rate of only 2.8\%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.
发表机构
- Fudan University(复旦大学)
- Hunyuan, Tencent(腾讯混元)
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。