发表机构
Indian Institute of Science; Hugging Face(印度科学研究所; 拥抱脸)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对文本到图像生成模型评估难题,提出DynEval动态评估框架,通过构建两个数据集,利用课程学习策略微调评估器,在多基准测试中与人类判断相关性更高,可对多个T2I模型进行细粒度分析。
AI 中文摘要
文本到图像(T2I)生成技术的进展使模型能生成高度逼真的图像,但可靠评估其输出仍具挑战。现有自动评估器难以捕捉细微失败模式。本文引入DynEval动态评估框架,用于联合评估T2I模型的文本到图像对齐和图像质量。构建了GenDB和DynEvalInstruct两个数据集,通过课程学习策略对紧凑型评估器进行全量微调,得到DynEval-2B和DynEval-4B。在11个基准测试中,该评估器与人类判断的整体相关性更高,还能对36个T2I模型进行细粒度分析。
英文摘要
Recent advances in text-to-image (T2I) generation have led to models capable of producing highly realistic images. Yet, reliably evaluating their outputs remains challenging, especially at scale. Existing automatic evaluators, often relying on a static prompt set, struggle to capture subtle failure modes such as partial prompt misalignment, compositional errors, or visually plausible but semantically incorrect generations. In this work, we introduce DynEval, a Dynamic Evaluation framework designed to jointly assess text-to-image alignment and image quality of T2I models. To support scalable training beyond limited human-annotated data, we construct two large datasets. First, we build GenDB, a collection of 500K prompt-image pairs generated from human-written prompts drawn from DiffusionDB using a tiered prompt-model generation strategy. Second, building upon GenDB, we construct DynEvalInstruct, a 250K instruction dataset comprising prompt-image-response triplets distilled from a structured evaluation pipeline that decomposes evaluation into text-image alignment and visual quality reasoning. Using this dataset, we perform full fine-tuning of a compact evaluator through a curriculum learning strategy to effectively distill the superior evaluation capabilities of a larger teacher vision-language model, resulting in DynEval-2B and DynEval-4B. In extensive comparisons against existing evaluators across 11 benchmarks, our evaluator achieves a higher overall correlation with human judgments. Furthermore, it provides fine-grained analysis of the capabilities and failure modes of 36 T2I models across 42 subcategories and 9 semantic dimensions.
CommentsAccepted at ECCV 2026. Project page: https://vcl-iisc.github.io/dyneval/