发表机构
Technical University of Munich; Helmholtz Munich; Munich Center for Machine Learning(慕尼黑工业大学; 亥姆霍兹慕尼黑中心; 慕尼黑机器学习中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文生图模型忽视否定约束评估的问题,提出含4,800提示词的极性基准NegT2IBench,通过检测器评分揭示多数模型在否定陈述上表现更差,并直接衡量模型抑制禁止内容的能力。
AI 中文摘要
文生图(T2I)模型通过衡量请求内容是否出现的基准来评判,但这些基准在很大程度上忽视了满足否定约束的互补能力,例如生成“一个非红色的杯子”。衡量否定带来了肯定式基准所没有的挑战,需要仔细设计提示词和评估方案。我们引入了NegT2IBench,一个包含4,800个提示词的基准,涵盖两种属性类型和四种关系类别。提示词按极性组织:必须成立的肯定陈述数量与必须不成立的否定陈述数量,各自取值范围为0到2。独立变化这两个维度,可将否定的影响与提示词复杂度的影响分离开来。我们的基于检测器的评分是可复现、可审计的,并能精确定位哪项要求失败。在包含600张图像、每张由三位标注者标注的测试集上,它与人类的一致性接近规模高达30倍的视觉-语言评判模型,同时仅使用其GPU内存的一小部分。在11个T2I模型和211,200张图像上,九个模型在单个否定陈述上的得分低于单个肯定陈述。逐陈述评分揭示,损失在颜色上最大,在邻近关系上几乎为零,且41.5%的失败陈述恰好渲染了提示词所禁止的内容。渲染提示词所要求的内容与抑制其所禁止的内容是两种不同的能力,而聚合的组合评分无法区分它们。NegT2IBench直接衡量后者,为诊断否定失败和开发克服方法提供了受控的测试平台。
英文摘要
Text-to-image (T2I) models are judged by benchmarks that measure whether requested content appears, but these benchmarks largely overlook the complementary ability to satisfy negated constraints, for example, generating "a non-red cup." Measuring negation raises challenges not faced by affirmation-based benchmarks and requires careful prompt and evaluation design. We introduce NegT2IBench, a benchmark of 4,800 prompts covering two attribute types and four relation categories. Prompts are organized by polarity: the number of positive statements that must hold and negated statements that must not, each ranging from 0 to 2. Varying the two independently separates the effect of negation from the effect of prompt complexity. Our detector-based scoring is reproducible, auditable, and pinpoints which requirement failed. On 600 images with three-annotator labels, it agrees with humans as closely as vision-language judges up to 30x larger, while using only a fraction of their GPU memory. Across eleven T2I models and 211,200 images, nine score lower on a single negated statement than on a single positive one. Per-statement scoring reveals that the loss is largest for color and near zero for proximity, and that 41.5% of failed statements render exactly what the prompt forbids. Rendering what a prompt asks for and withholding what it forbids are distinct capabilities that an aggregate compositional score cannot distinguish. NegT2IBench measures the latter directly, providing a controlled testbed for diagnosing negation failures and developing methods to overcome them.
Comments*Equal contribution