arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成式图像模型中的持久身份保留:基准与评估系统

Persistent Identity Preservation in Generative Image Models: A Benchmark and Evaluation System

Mengwei Ren, Xuaner Zhang, Zhihao Xia

arXiv 2609.04151首次发表:更新:

发表机构

Phota Labs Research(Phota实验室研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对生成式图像模型的身份保留问题构建基准,对比不同身份表示范式,发现PHOTA IDENTITY等持久身份方案可有效缓解身份退化,提升多场景下的身份保真度。

AI 中文摘要

当前生成式图像模型已能生成高质量图像、遵循复杂指令并支持精确编辑,但仍难以保留所描绘对象的身份。在生成或编辑特定对象的图像时,若姿态、表情、外观、视角或周围场景发生变化,身份可能会出现偏移。现有对象驱动方法对于身份的表示位置存在根本不同的选择:通过输入上下文(GPT-Image-2、NB2)、可训练的对象特定模型参数(LoRA),或通过可在生成与编辑中复用的持久身份层(PHOTA IDENTITY)。我们针对对象驱动生成、编辑、修复及多对象设置,对这些范式进行了系统基准测试,所设计的任务会逐步增加身份保留的压力。我们的结果表明,身份保留仍是当前生成式基础模型的显著局限:强大的图像质量与指令遵循能力并不必然意味着高身份保真度,且在迭代编辑、小对象尺度、严重图像退化及多对象组合场景下,身份退化会更为明显。持久身份可大幅降低生成、编辑与修复过程中的这种退化,在应用于不同基础模型时能持续提升身份保留效果,同时保持相当的指令遵循能力与感知图像质量。这些结果表明,身份并非简单地从日益强大的生成式模型中自然产生,而是可被表示为独立于底层生成式模型组合的持久对象知识。

英文摘要

Generative image models can now produce high-quality images, follow complex instructions, and support precise edits, but they still struggle to preserve who or what is being depicted. When generating or editing images of a specific subject, identity may drift as the pose, expression, appearance, viewpoint, or surrounding scene changes. Existing subject-driven methods make fundamentally different choices about where identity is represented: through the input context (GPT-Image-2, NB2), as trainable subject-specific model parameters (LoRA), or as a persistent identity layer (PHOTA IDENTITY) reusable across generations and edits. We systematically benchmark these paradigms across subject-driven generation, editing, restoration, and multi-subject settings, with tasks designed to increasingly stress identity preservation. Our results show that identity preservation remains a distinct limitation of current generative foundation models: strong image quality and instruction following do not necessarily imply strong identity fidelity, and identity degradation becomes more pronounced under iterative edits, small subject scales, severe image degradation, and multi-subject composition. Persistent identity substantially reduces this degradation across generation, editing, and restoration, consistently improving identity preservation when applied to different foundation models while maintaining comparable instruction adherence and perceptual image quality. These results suggest that identity does not simply emerge from increasingly capable generative models, but can instead be represented as persistent subject knowledge that is composed independently with the underlying generative model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑