arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DynEval:对自然环境下的文本到图像生成模型进行全面评估

DynEval: Holistic Evaluations of T2I Generative Models in the Wild

Shyam Marjit, Dheeraj Baiju, Anuj Shikarkhane, Akhil Sakthieswaran, Sayak Paul, Anirban Chakraborty

arXiv 2607.11199首次发表:更新:

发表机构

Indian Institute of Science; Hugging Face(印度科学研究所; 拥抱脸)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对文本到图像生成模型评估难题,提出DynEval动态评估框架,通过构建两个数据集,利用课程学习策略微调评估器,在多基准测试中与人类判断相关性更高,可对多个T2I模型进行细粒度分析。

AI 中文摘要

文本到图像(T2I)生成技术的进展使模型能生成高度逼真的图像,但可靠评估其输出仍具挑战。现有自动评估器难以捕捉细微失败模式。本文引入DynEval动态评估框架,用于联合评估T2I模型的文本到图像对齐和图像质量。构建了GenDB和DynEvalInstruct两个数据集,通过课程学习策略对紧凑型评估器进行全量微调,得到DynEval-2B和DynEval-4B。在11个基准测试中,该评估器与人类判断的整体相关性更高,还能对36个T2I模型进行细粒度分析。

英文摘要

Recent advances in text-to-image (T2I) generation have led to models capable of producing highly realistic images. Yet, reliably evaluating their outputs remains challenging, especially at scale. Existing automatic evaluators, often relying on a static prompt set, struggle to capture subtle failure modes such as partial prompt misalignment, compositional errors, or visually plausible but semantically incorrect generations. In this work, we introduce DynEval, a Dynamic Evaluation framework designed to jointly assess text-to-image alignment and image quality of T2I models. To support scalable training beyond limited human-annotated data, we construct two large datasets. First, we build GenDB, a collection of 500K prompt-image pairs generated from human-written prompts drawn from DiffusionDB using a tiered prompt-model generation strategy. Second, building upon GenDB, we construct DynEvalInstruct, a 250K instruction dataset comprising prompt-image-response triplets distilled from a structured evaluation pipeline that decomposes evaluation into text-image alignment and visual quality reasoning. Using this dataset, we perform full fine-tuning of a compact evaluator through a curriculum learning strategy to effectively distill the superior evaluation capabilities of a larger teacher vision-language model, resulting in DynEval-2B and DynEval-4B. In extensive comparisons against existing evaluators across 11 benchmarks, our evaluator achieves a higher overall correlation with human judgments. Furthermore, it provides fine-grained analysis of the capabilities and failure modes of 36 T2I models across 42 subcategories and 9 semantic dimensions.

CommentsAccepted at ECCV 2026. Project page: https://vcl-iisc.github.io/dyneval/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑