发表机构
Khalifa University; Queen Mary University of London; The Chinese University of Hong Kong; The University of Western Australia; The University of Melbourne; University of Central Florida(哈利法大学; 伦敦玛丽女王大学; 香港中文大学; 西澳大学; 墨尔本大学; 中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对多语言文本到图像生成的跨语言一致性局限,构建含10种语言、3.3万条提示词的LingT2I基准,经分析揭示语言相关生成规律,为开发更鲁棒的多语言T2I模型奠定基础。
AI 中文摘要
文本到图像(T2I)生成近年来取得了显著进展,但现有研究大多仅关注英语场景,跨语言性能差距及语言特定效应尚未得到充分探索。为填补这一空白,本文引入LingT2I基准,覆盖10种常用语言,包含3.3万条提示词,旨在评估内容生成与文本渲染两方面的跨语言效应。基于该基准,本文开展了全面的跨语言分析,揭示了各评估维度下的语言不平等及语言依赖权衡;除定量评估外,还发现一系列语言依赖的生成模式,凸显语言因素及其对应文化背景如何系统性影响模型输出。本文的基准与分析为研究T2I生成的跨语言行为提供了基础,有助于开发更鲁棒、更具包容性的模型,代码与数据集可在指定URL获取。
英文摘要
Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual performance gaps and language-specific effects insufficiently explored. To fill this gap, we introduce LingT2I, a benchmark covering 10 widely used languages with 33K prompts, designed to evaluate cross-lingual effects in both content generation and text rendering. Building on this benchmark, we conduct a comprehensive cross-lingual analysis, uncovering linguistic inequality and language-dependent trade-offs across evaluation dimensions. Beyond quantitative evaluation, we further reveal a range of language-dependent generation patterns, highlighting how linguistic factors and their corresponding cultural contexts systematically impact model outputs. Our benchmark and analysis provide a foundation for studying cross-lingual behavior in T2I generation and facilitate the development of more robust and inclusive models. Code and dataset are available at https://github.com/RISys-Lab/LingT2I.
CommentsAccepted to ACM MM 2026