发表机构
The University of Tokyo; Nara Institute of Science and Technology; Chungnam National University; Institute of Science Tokyo(东京大学; 奈良科学技术研究所; 忠南国立大学; 东京科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本到图像模型常混淆明喻喻体与本体的问题,提出含受控数据集、YOLO指标及Diffusion Lens分析的评估框架,实验发现模型存在字面化失败模式并讨论了缓解策略。
AI 中文摘要
明喻为文本提示中描述视觉特征提供了简洁且富有表现力的方式。近期的文本到图像模型(t2i模型)能够根据明喻提示生成视觉上有吸引力的输出,但即使是前沿模型也经常错误理解隐喻喻体,将其与本体混淆。这些系统性失败揭示了文本到图像模型中比喻语言与对象级视觉定位之间的差距。为研究该问题,我们提出了一个可扩展的明喻理解评估框架,该框架包含三部分:(1)受控明喻数据集,其中隐喻喻体选自一组预定义的可检测对象类别,并与多样模板组合;(2)基于YOLO(You Only Look Once)检测的自动定位指标;(3)使用Diffusion Lens进行的文本编码器层分析,以追踪生成过程中隐喻喻体的出现情况。对不同架构的文本到图像模型开展的实验揭示了一致的字面化失败模式,我们还进一步讨论了改善文本到图像模型中明喻定位的潜在缓解策略。
英文摘要
Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce visually compelling outputs from simile prompts, yet even frontier models frequently misinterpret the metaphorical vehicle and confuse it with the object. These systematic failures reveal a gap between figurative language and object-level visual grounding in t2i models. To investigate this issue, we propose a scalable evaluation framework for simile understanding. Our framework includes (1) a controlled simile dataset in which metaphorical vehicles are drawn from a predefined set of object-detectable categories and combined with diverse templates, (2) automatic grounding metrics based on YOLO (You Only Look Once) detection, and (3) text encoder layer analysis using Diffusion Lens to track how metaphorical vehicles emerge during generation. Experiments across architecturally diverse t2i models reveal consistent literalization failure patterns. We further discuss potential mitigation strategies for improving simile grounding in t2i models.
CommentsAccepted as a full paper at ACM Multimedia 2026