arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于图像-描述提示的前沿文本到图像模型基准测试

Benchmarking Frontier Text-to-Image Models on Image-Description Prompts

Sajjad Abdoli, Ghassan Al-Sumaidaee, Ahmed Rashad

arXiv 2608.14976首次发表:更新:

发表机构

Perle(珀尔)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对组合要求高的图像-描述提示,评估了四个前沿文本到图像模型的性能,发现 Gemini 3 Pro Image 表现最优,领先系统的主要问题是对象计数错误和几何伪影。

AI 中文摘要

文本到图像模型通常在平均案例提示上进行报告,这低估了系统在涉及精确对象计数、多对象属性绑定、清晰嵌入文本和显式空间约束等组合要求高的请求上的性能差距。我们评估了四个生产级文本到图像系统:Hunyuan 3.0、Gemini 3 Pro Image(“Nano Banana Pro”)、Black Forest Labs FLUX.2 和 Ideogram 3.0。评估使用了从 this http URL 样本数据集(DSD)中通过对全部语料库进行自动化复杂度评分筛选出的 48 个最难提示。每张生成图像都使用独立评判标准进行评分:GPT-5.4-Pro 制定了原子化、加权、互斥且完全穷尽(MECE)的评估标准,而 Gemini 3.1 Pro Preview 则独立确定每个标准是否满足。Gemini 3 Pro Image 以 84.8/100 的分数排名第一,略高于得分为 82.3/100 的 FLUX.2;Ideogram 3.0 和 Hunyuan 3.0 的得分分别为 65.7/100 和 63.3/100。失败分析显示,领先系统主要因对象计数错误和几何伪影失分,而落后系统更常生成乱码文本,Ideogram 3.0 还频繁遗漏请求的元素。完整的逐样本标准、分数和失败注释可向作者索取。

英文摘要

Text-to-image models are typically reported on average-case prompts, which understates the gap between systems on compositionally demanding requests involving precise object counts, multi-object attribute binding, legible embedded text, and explicit spatial constraints. We evaluate four production text-to-image systems: Hunyuan 3.0, Gemini 3 Pro Image ("Nano Banana Pro"), Black Forest Labs FLUX.2, and Ideogram 3.0. The evaluation uses the 48 hardest prompts drawn from the DataSeeds.AI Sample Dataset (DSD), selected through an automated complexity-scoring pass over the full corpus. Every generated image is graded using an independent-judge rubric. GPT-5.4-Pro authors an atomic, weighted, mutually exclusive and collectively exhaustive (MECE) evaluation rubric, while Gemini 3.1 Pro Preview independently determines whether each criterion is satisfied. Gemini 3 Pro Image ranks first with a score of 84.8/100, narrowly ahead of FLUX.2 at 82.3/100. Ideogram 3.0 and Hunyuan 3.0 score 65.7/100 and 63.3/100, respectively. Failure analysis shows that the leading systems primarily lose points through object miscounting and geometric artifacts, whereas the trailing systems more frequently produce garbled text. Ideogram 3.0 also frequently omits requested elements. Full per-sample rubrics, scores, and failure annotations are available from the authors upon request.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑