arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CopyCat:在数秒内提升主体到图像模型中的细粒度主体一致性

CopyCat: Improving Fine-Grained Subject Consistency in Subject-to-Image Models within Seconds

Peng Zheng, Ruiqi Liu, Rui Ma, Zuxuan Wu

arXiv 2608.00674首次发表:更新:

AI 中文总结

本研究提出轻量级模型精调框架CopyCat,通过附加FCLoRA在数秒内提升主体到图像模型的细粒度主体一致性,经DreamBench等实验验证,可适配多样未见过的主体与提示并取得持续效果。

AI 中文摘要

近期的主体到图像模型在个性化图像生成方面已取得令人瞩目的进展,但仍难以保留细粒度的主体特定细节。一个主要原因是缺乏高质量的细粒度身份监督:真实配对数据的收集成本高昂,而合成的训练配对通常仅保留主体的粗粒度外观,无法捕捉细微的主体特定细节。在本研究中,我们提出了CopyCat,这是一个轻量级模型精调框架,可在仅数秒内提升细粒度主体一致性。CopyCat通过附加轻量级细粒度一致性LoRA(FCLoRA)并使用单一代理图像对其进行优化,对预训练的主体到图像模型执行一次性精调,该代理图像同时作为条件图像和重建目标。这种精确的自重建目标大幅简化了优化任务,使模型能在仅数秒内实现有效的细粒度精调。该精调仅执行一次;生成的模型可直接应用于多样的未见过的参考主体和提示,无需进一步的主体特定优化。我们还重新审视了双流扩散Transformer中的主体到图像LoRA训练,发现仅适配视觉流可持续提升主体一致性。在DreamBench和XVerseBench上的大量实验表明,在单主体和多主体设置下,CopyCat均能在代表性主体到图像模型中实现细粒度主体一致性的持续提升。

英文摘要

Recent subject-to-image models have achieved impressive progress in personalized image generation, yet they still struggle to preserve fine-grained subject-specific details. A major reason is the lack of high-quality fine-grained identity supervision: real paired data are expensive to collect, while synthesized training pairs often preserve only coarse subject appearance and fail to capture subtle subject-specific details. In this work, we propose CopyCat, a lightweight model-refinement framework that improves fine-grained subject consistency within only a few seconds. CopyCat performs a one-time refinement of a pretrained subject-to-image model by attaching a lightweight Fine-grained Consistency LoRA (FCLoRA) and optimizing it using a single proxy image, which is used as both the conditioning image and the reconstruction target. This exact self-reconstruction objective substantially simplifies the optimization task, enabling effective fine-grained refinement within only a few seconds. The refinement is performed only once; the resulting model can be directly applied to diverse unseen reference subjects and prompts without further subject-specific optimization. We further revisit subject-to-image LoRA training in double-stream diffusion transformers and find that adapting only the visual stream consistently improves subject consistency. Extensive experiments on DreamBench and XVerseBench demonstrate consistent improvements in fine-grained subject consistency across representative subject-to-image models under both single- and multi-subject settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑