arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11885cs.CVcs.GR

文本到图像个性化模型中的潜在身份调整

Latent-Identity Tuning in Text-to-Image Personalization Models

Daniel Garibi, Ronen Kamenetsky, Hadar Averbuch-Elor, Daniel Cohen-Or, Or Patashnik

首次发表
浏览论文内容

中文总结 AI 辅助

研究文本到图像个性化模型中细粒度身份调整问题,利用预训练冻结编码器潜在空间,无需额外训练,通过揭示潜在语义方向实现局部、细粒度且语义连贯的面部编辑,经实验验证有效。

中文摘要 AI 辅助

生成和编辑人脸需要高精度,因为即使是细微修改也可能显著改变人物身份。当前基于通用文本到图像模型的个性化和编辑方法往往缺乏细粒度面部编辑所需的精度。我们提出一种文本到图像个性化模型中的细粒度身份调整方法。与在给定图像上操作的标准图像编辑不同,身份调整修改特定身份的潜在表示,能生成一致描绘相同编辑身份的多样图像。为实现细粒度潜在身份调整,我们探索预训练、冻结编码器的潜在空间。该方法无需额外训练,利用冻结编码器现有架构揭示潜在语义方向。此空间由一组潜在令牌组成,它们在捕捉身份不同方面发挥不同作用,常对应特定空间或语义面部区域。我们表明可在此空间及所选令牌定义的子空间内识别有意义的方向,实现局部、细粒度且语义连贯的编辑。通过定性和定量实验验证了该方法,展示了多样的局部面部编辑,同时保持跨图像身份一致性。

英文摘要

Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a subject's perceived identity. Current personalization and editing methods built on general-purpose text-to-image models, however, often lack the precision required for fine-grained facial edits. We present a method for fine-grained identity tuning in text-to-image personalization models. Unlike standard image editing, which operates on a given image, identity tuning modifies the latent representation of a specific identity, enabling the generation of diverse images that consistently depict the same edited identity. To enable fine-grained latent identity tuning, we explore the latent space of a pre-trained, frozen encoder for text-to-image personalization. Our approach requires no additional training. Instead, it leverages the existing architecture of a frozen encoder to uncover latent semantic directions. This space consists of a set of latent tokens that play distinct roles in capturing different aspects of an identity and often correspond to specific spatial or semantic facial regions. We show that meaningful directions can be identified within this space and within subspaces defined by selected tokens, enabling localized, fine-grained, and semantically coherent edits. We validate our approach through qualitative and quantitative experiments that demonstrate diverse localized facial edits while preserving cross-image identity consistency. Project page at: https://garibida.github.io/IdentityTuning/

发表机构

  • Tel Aviv University(特拉维夫大学)
  • Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑