发表机构
Harbin Institute of Technology, Shenzhen; Sun Yat-sen University; South China University of Technology; Guangdong University of Technology(哈尔滨工业大学(深圳); 中山大学; 华南理工大学; 广东工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本到图像生成模型对提示制定敏感的问题,提出PRISM框架,通过基于图像的自我奖励机制,利用结构化视觉诊断和多任务监督微调等方法,提高图像质量和语义对齐,并提供可解释反馈。
AI 中文摘要
文本到图像生成模型可根据自然语言描述合成高质量图像,但其性能对提示制定高度敏感。现有提示优化方法主要依赖文本重写、提示扩展或外部奖励信号,对基于图像的诊断支持有限。本文提出PRISM,通过基于图像的自我奖励机制进行提示优化。PRISM通过结构化视觉诊断解释生成图像,并根据语义一致性、美学质量和人类偏好对齐进行评分,闭合提示-图像反馈回路。它先通过多任务监督微调初始化统一的VLM,然后通过混合理想点和切比雪夫奖励的自我奖励优化改进提示策略。大量实验表明,PRISM提高了整体图像质量和细粒度语义对齐,同时为有针对性的提示优化提供可解释反馈。代码可在指定网址获取。
英文摘要
Text-to-image generation models can synthesize high-quality images from natural language descriptions, but their performance remains highly sensitive to prompt formulation. Existing prompt optimization methods mainly rely on text-side rewriting, prompt expansion, or external reward signals, offering limited image-grounded diagnosis and weak support for learning reusable optimisation policies. In this paper, we propose PRISM, a Prompt Refinement framework via Image-grounded Self-rewarding Mechanism. PRISM closes the prompt-image-feedback loop by interpreting generated images with structured visual diagnosis and scoring them along semantic consistency, aesthetic quality, and human preference alignment. It first initializes a unified VLM through multi-task supervised fine-tuning, and then improves the prompt policy via self-rewarding optimization with a hybrid ideal-point and Chebyshev reward. Extensive experiments show that PRISM improves holistic image quality and fine-grained semantic alignment, while providing interpretable feedback for targeted prompt refinement. The code is available at https://anonymous.4open.science/r/PRISM-FF81.
Comments18 pages, 5 figures