arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一劳永逸?用于文本到图像提示优化的类型感知修复分配

One Rewrite to Fix Them All? Type-Aware Repair Allocation for Text-to-Image Prompt Optimization

Haoyue Liu, Xiaoyu Ma, Ye Chen, Shuguang Cui, Xiaoying Tang

arXiv 2607.18724首次发表:更新:

发表机构

School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen; XJTU-POLIMI Joint School, Xi’an Jiaotong University; Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen)(香港中文大学(深圳)理工学院; 西安交通大学-米兰理工大学联合学院; 深圳未来智能网络研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究文本到图像提示优化问题,提出将语义提示优化制定为原子修复分配的方法,在无需训练的TARA框架中实例化,实验表明该方法在多个基准生成器单元中语义准确性最佳,且运行快、图像质量好。

AI 中文摘要

文本到图像(T2I)生成器常常无法忠实地遵循提示,出现数量错误、属性交换、关系模糊和文本难以辨认等问题。提示优化通过重写用户提示来修复此类故障,无需重新训练生成器,已取得了有前景的成果。然而,现有优化器将不同类型的故障统一进行提示扩展,尽管每种故障需要不同的修复语言。我们将语义提示优化制定为原子修复分配:每个失败的命题在最终的局部约束被编译成一个可执行提示之前,被路由到一个类型条件修复操作符。我们在无需训练的类型感知修复分配(TARA)框架中实例化了这个公式,该框架分离了诊断、分配、编译和一个语义修复门,这是一个对恰好一个规定修复的接受或恢复控制器,可防止语义回归。在四个冻结生成器上对DSG和TIFA进行的广泛实验表明,TARA在所有八个基准生成器单元中实现了最佳语义准确性,在DSG和TIFA上分别比VisualPrompter提高了5.6和2.6个百分点,同时保持图像质量,并且在我们匹配的本地设置中运行最快,每个提示16.0秒,而VisualPrompter为20.0秒。

英文摘要

Text-to-image (T2I) generators often fail to follow their prompts faithfully, producing wrong counts, swapped attributes, ambiguous relations, and illegible text. Prompt optimization repairs such failures by rewriting the user prompt, requiring no generator retraining, and has yielded promising results. However, existing optimizers absorb heterogeneous failures into one uniform prompt expansion, even though each calls for different repair language. We formulate semantic prompt optimization as atomic repair allocation: each failed proposition is routed to a type-conditioned repair operator before the resulting local constraints are compiled into one executable prompt. We instantiate this formulation in the training-free Type-Aware Repair Allocation (TARA) framework, which separates diagnosis, allocation, compilation, and a semantic repair gate, an accept-or-revert controller over exactly one prescribed repair that prevents semantic regressions. Extensive experiments on DSG and TIFA across four frozen generators demonstrate that TARA achieves the best semantic accuracy in all eight benchmark-generator cells, improving over VisualPrompter by 5.6 and 2.6 points on DSG and TIFA, respectively, while maintaining image quality and running fastest in our matched local setting at 16.0 seconds versus 20.0 seconds per prompt.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑