arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UltraText Bench:评估图像生成中视觉文本渲染的全面双语基准

UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation

Deyuan Liu, Yihao Hu, Jingxuan Zhang, Xingying Li, Jun Xie, Jiacheng Liu, Jungang Li, Yu Huang, Xuanyi Liu, Yue Ding, Zecheng Wang, Lei Zhao, Mingda Wang, Zhenglin Cheng, Peng Sun, Tao Lin

arXiv 2610.09823首次发表:更新:

发表机构

Westlake University; Ant Group; Zhejiang University; Shanghai Innovation Institute; HKUST; CASIA; Wechat AI; MBZUAI; CityU; Peking University(西湖大学; 蚂蚁集团; 浙江大学; 上海创新研究院; 香港科技大学; 中国科学院自动化研究所; 微信人工智能; 穆罕默德·本·扎耶德人工智能大学; 香港城市大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

UltraText Bench是一个双语基准,用于评估图像生成中密集视觉文本的渲染,通过432个提示词和Q-Judger模型评估文本保真度、清晰度、空间和场景质量,发现不同模型在不同维度上的性能差异。

AI 中文摘要

密集视觉文本要求图像生成器在多个区域中再现长字符串,并保证正确的位置和可读性。随着短字符串渲染的改进,评估必须测试在更具挑战性的场景中的持续性能。我们引入了UltraText Bench,这是一个用于提示词驱动的密集视觉文本生成的双语基准。它包含432个提示词,涵盖24个真实世界场景类别和三个难度级别,中英文各占一半。每个经过人工审核的提示词提供四到十二个文本区域的精确字符串,并配有关于其内容、位置和视觉属性的结构化参考。我们使用Q-Judger视觉语言模型根据完整参考评估每张图像,报告文本保真度、文本清晰度、空间质量和场景质量。在24种模型配置中,这些维度揭示了不同的优势:Z-Image-Turbo在报告设置下比Z-Image-Base获得3.81清晰度点,同时损失14.76保真度点。性能也随工作量变化;Qwen-Image-2512的英文综合得分从L1的86.50降至L3的42.86。十名参与者参与了自动评分的人工评估。代码库:此https URL。

英文摘要

Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation of dense visual text. It contains 432 prompts spanning 24 real-world scene categories and three difficulty levels, split equally between English and Chinese. Each human-reviewed prompt supplies exact strings for four to twelve text regions, paired with structured references for their content, placement, and visual attributes. We use the Q-Judger vision-language model to assess each image against the complete reference, reporting text fidelity, text clarity, spatial quality, and scene quality. Across 24 model configurations, these dimensions reveal different strengths: Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points under the reported settings. Performance also varies with workload; Qwen-Image-2512's English composite falls from 86.50 at L1 to 42.86 at L3. Ten participants took part in human evaluation of the automatic scores. Repository: https://github.com/LINs-lab/UltraText_Bench.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑