Imag-Eval:一种基于语言的可解释文本到图像指令遵循评估框架
Imag-Eval: a language-grounded framework for interpretable Text-to-Image instruction following evaluation
查看机构详情
- Talan Research and Innovation Center(塔兰研究与创新中心)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对现有T2I模型评估的局限,提出Imag-Eval基准,通过独立调整实例数与规则组合区分语言复杂性与组合难度,经实验表明结构化技能的组合难度由规则数及绑定方式决定。
中文摘要 AI 辅助
文本到图像(Text-to-Image,T2I)模型近期在视觉保真度方面取得了令人瞩目的成果,但其评估仍受限于难以解释且诊断性不足的基准。现有的基于技能的评估往往忽略了对可用性影响重大但不在标准分类内的关键失败模式,例如因缺失部分或物理上不合理的配置(如漂浮物体)导致的全局不一致。此外,提示难度通常仅沿单一维度控制,要么是提示长度,要么是要生成的元素数量。为解决这些局限,我们引入了Imag-Eval,这是一个受控基准,旨在评估T2I模型如何将组合式自然语言指令落地为视觉输出。与之前将表面语言复杂性与组合难度混为一谈的工作不同,Imag-Eval通过独立调整实例数量和约束(规则)的组合来明确区分这些因素,同时避免错误传播。该设计支持对跨模态指令遵循失败的具体位置进行细粒度且可解释的分析。我们的基准包含1140个提示和8842条组合规则,并在多个最先进的模型上对其进行了评估。结合对来自同期基准的2000多个提示的额外研究,我们的结果表明,对于结构化技能,组合难度主要由已落地规则的数量及其与实例的绑定决定,而非仅由提示长度决定。
英文摘要
Text-to-Image (T2I) models have recently achieved impressive visual fidelity, yet their evaluation remains constrained by benchmarks that are often difficult to interpret and insufficiently diagnostic. Existing skill-based evaluations tend to overlook critical failure modes that strongly impact usability but fall outside standard taxonomies, such as global incoherence arising from missing parts or physically implausible configurations (e.g., floating objects). In addition, prompt difficulty is typically controlled along a single dimension; either prompt length or the number of elements to generate. To address these limitations, we introduce Imag-Eval, a controlled benchmark designed to assess how T2I models ground compositional natural-language instructions into visual outputs. Unlike prior work that conflates surface linguistic complexity with compositional difficulty, Imag-Eval explicitly seeks to disentangles these factors by independently varying both the number of instances and the combination of constraints (rules), while avoiding error propagation. This design enables fine-grained and interpretable analysis of where cross-modal instruction following fails. Our benchmark comprises 1,140 prompts and 8,842 combined rules, and we evaluate it on several state-of-the-art models. Complementing this analysis with an additional study of over 2,000 prompts from a concurrent benchmark, our results suggest that, for structured skills, compositional difficulty is primarily governed by the number of grounded rules and their binding to instances,, rather than by prompt length alone.