arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有推理与反思的个性化图像生成

Personalized Image Generation with Reasoning and Reflection

Bo Ni, Ngoc N. Tran, Qinwen Ge, Franck Dernoncourt, Seunghyun Yoon, Samyadeep Basu, Sungchul Kim, Puneet Mathur, Nedim Lipka, Tong Yu, Yu Wang, Ryan A. Rossi, Tyler Derr

arXiv 2610.00737首次发表:更新:

发表机构

Vanderbilt University; Adobe Systems; University of Maryland, College Park; University of Georgia(范德比尔特大学; Adobe系统公司; 马里兰大学学院公园分校; 佐治亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对个性化图像生成,提出首个基于用户历史的统一基准,涵盖场景与创意生成两任务,并推出PEARL模型,通过推理-反思循环优化,平均提升个性化指标15%。

AI 中文摘要

个性化图像生成一直局限于从精选的视觉示例中进行条件合成,而非捕捉用户是谁。然而,在实践中,用户的个人背景要丰富得多,包括随时间积累的评论、帖子、图像、标题和元数据。一个真正的个性化生成器应利用这些历史信息来生成与用户生活方式和审美偏好一致的图像。为此,我们引入了首个基于用户历史进行个性化图像生成的统一基准。该基准包含两个互补的任务和一个多轴评估协议,该协议评估目标保真度、视觉质量、用户可区分性、与用户历史的语义一致性以及任务特定效用。基于真实的电子商务和社交媒体场景,该基准包括:(1) 个性化场景生成,将给定对象放置在反映用户偏好和生活方式的场景中,其动机源于个性化产品展示;(2) 个性化创意生成,在指定主题上生成新颖图像,该图像忠实于用户的审美和视觉身份,其动机源于社交媒体内容创作。我们进一步提出了PEARL,它将多模态推理器与冻结的图像生成器在交错推理-反思循环中耦合,并通过差分数据奖励进行优化。在这两个任务中,PEARL均优于强基线,在个性化指标上平均提升了15%。

英文摘要

Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user is. In practice, however, a user's personal context is much richer, comprising reviews, posts, images, captions, and metadata accumulated over time. A truly personalized generator should leverage this history to produce images aligned with the user's lifestyle and aesthetic preferences. To this end, we introduce the first unified benchmark for personalized image generation from user histories. The benchmark comprises two complementary tasks and a multi-axis evaluation protocol that assesses target fidelity, visual quality, user distinguishability, semantic alignment with the user's history, and task-specific utility. Grounded in real-world e-commerce and social media settings, the benchmark includes: (1) Personalized Scene Generation, which places a given object in a scene that reflects a user's preferences and lifestyle, motivated by personalized product presentation; and (2) Personalized Creative Generation, which generates a novel image on a specified topic that is faithful to a user's aesthetic and visual identity, motivated by social media content creation. We further propose PEARL, which couples a multimodal reasoner with a frozen image generator in an interleaved reason-reflect loop optimized with differential data reward. Across both tasks, PEARL outperforms strong baselines, achieving an average improvement of 15% across personalization metrics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑