arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02502cs.CVcs.AI

融合概念:文本到图像模型中视觉隐喻生成的基准测试

Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models

Chuer Chen, Zichen Wang, Yi He, Zhengxi Yu, Nan Cao

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对文本到图像(T2I)模型视觉隐喻生成能力未被充分研究的问题,推出首个基准VMetaphor-Bench,结合混合框架评估11个代表性T2I模型,发现其在隐喻关键方面存在不足,指明该领域为未来研究前沿。

中文摘要 AI 辅助

文本到图像(T2I)模型在忠实渲染指定对象和属性方面已取得显著成功,但它们生成视觉隐喻(结合两个不同领域元素以传达抽象思想的图像)的能力仍未得到充分研究。为填补这一空白,我们推出VMetaphor-Bench,首个用于评估T2I模型视觉隐喻生成的基准。它包含1500个从现实世界创意图像中精选的视觉隐喻,分为三个级别和十个类别,每个样本配有两个不同特异性的提示。评估时,我们在多模态大模型(MLLM)作为评判者的范式下开发了混合框架,结合了基于多项选择题(MCQ)的协议(涵盖四个隐喻保真度级别,共9594个问题)和沿三个感知维度的基于维度的评分协议。对11个代表性T2I模型的广泛评估显示,即使是最强的专有模型也在构图结构和跨域映射(隐喻表达的关键方面)上表现不佳,凸显视觉隐喻生成是未来T2I研究的重要前沿领域。

英文摘要

Text-to-image (T2I) models have achieved remarkable success at faithfully rendering specified objects and attributes, yet their ability to produce visual metaphors, images that convey abstract ideas by combining elements from two distinct domains, remains largely unexamined. To bridge this gap, we introduce VMetaphor-Bench, the first benchmark for evaluating visual metaphor generation in T2I models. It comprises 1,500 visual metaphors curated from real-world creative imagery, organized into three levels and ten categories, with each sample paired with two prompts of differing specificity. For evaluation, we develop a hybrid framework within an MLLM-as-judge paradigm, combining a multiple-choice question (MCQ) based protocol of 9,594 questions across four levels of metaphorical fidelity with a dimension-based scoring protocol along three perceptual dimensions. Extensive evaluation of 11 representative T2I models reveals that even the strongest proprietary models struggle with compositional structuring and cross-domain mapping, key aspects of metaphorical expression, highlighting visual metaphor generation as an important frontier for future T2I research.

发表机构

  • Tongji University(同济大学)
  • Shanghai Innovation Institute(上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑