arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14226cs.CV

RankT2I:一种用于发现文本到图像模型中可解释且多样化语义的子模框架

RankT2I: A Submodular Framework for Discovering Interpretable and Diverse Semantics in Text-to-Image Models

Ritika Allada, Pinar Yanardag

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出无训练、模型无关的RankT2I框架,通过子模方法自动发现T2I模型的可编辑语义,性能优于现有方法,可高效获取多领域文本到图像编辑所需的多样化语义。

中文摘要 AI 辅助

近期文本到图像(T2I)模型的进展彻底革新了图像生成与编辑领域。然而,识别T2I模型可成功编辑的语义仍是一项具有挑战性的任务。现有多数方法需用户手动指定语义以修改特定图像,该过程耗时且常涉及大量试错。本文提出RankT2I,一种新颖的无训练、模型无关框架,可自动发现扩散模型及基于FLUX的模型中的可编辑语义。给定视觉领域,我们首先利用多模态视觉-语言模型收集大量候选语义,随后将语义发现建模为集合选择问题,采用子模目标识别相关、可编辑且多样化的语义。该方法可帮助用户在多个领域中高效识别文本到图像编辑模型的广泛语义,且性能优于现有方法。

英文摘要

Recent advances in text-to-image (T2I) models have revolutionized the field of image generation and editing. However, identifying semantics that a T2I model can successfully edit in an image continues to be a challenging task. Most existing approaches require users to manually specify semantics to modify a particular image, a time-consuming process that often involves extensive trial and error. In this paper, we present RankT2I, a novel, training-free, and model-agnostic framework that automates the discovery of editable semantics in diffusion and FLUX-based models. Given a visual domain, we first utilize a multimodal vision-language model to gather a broad set of candidate semantics. We then frame semantic discovery as a set selection problem and use a submodular objective to identify semantics that are relevant, editable, and diverse. Our method helps users efficiently identify a wide range of semantics for text-to-image editing models across several domains while outperforming existing methods.

发表机构

  • Virginia Tech(弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑