arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32188cs.CVcs.AI

存在并非忠实:文生图中的喻体侵入

Presence Is Not Faithfulness: Figurative Vehicle Intrusion in Text-to-Image Generation

Xiaoyu Ma, Chen Yang, Hao Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对文生图模型将喻体误渲染为可见物体的问题,提出VISTA基准和V-Score指标,并引入VISTA-Guard缓解策略,以提升比喻忠实度。

中文摘要 AI 辅助

文生图(TTI)模型日益能够根据自然语言提示生成高质量图像,然而比喻性语言暴露了一个失败:本应引导本体描绘的喻体,反而可能被渲染为可见物体。我们将此失败称为“喻体侵入”:侵入内容在文本上得到许可,但被赋予了错误的视觉角色,这表明视觉存在并不总是忠实,且以存在为导向的评估可能遗漏此类错误。为系统研究此现象,我们引入了“喻体侵入与语义本体评估”(VISTA),一个按比喻形式和映射机制组织的多语言比喻提示基准。我们进一步提出V-Score,一种诊断性问答指标,用于评估生成图像中角色感知的比喻忠实度。对近期高性能TTI模型的评估显示,喻体侵入在多种语言和比喻类别中持续存在。作为一种轻量级缓解措施,我们引入VISTA-Guard,它能部分减少喻体侵入,并为更比喻忠实的TTI生成提供了一条实用路径。所有资源将公开发布。

英文摘要

Text-to-image (TTI) models increasingly generate high-quality images from natural-language prompts, yet figurative language exposes a failure: a vehicle that should guide the depiction of a tenor may instead be rendered as a visible object. We call this failure Figurative Vehicle Intrusion: the intruding content is textually licensed, but it is assigned the wrong visual role, showing that visual presence is not always faithfulness and that presence-oriented evaluation can miss such errors. To study it systematically, we introduce Vehicle Intrusion and Semantic Tenor Assessment (VISTA), a multilingual benchmark of figurative prompts organized by Figurative Form and Mapping Mechanism. We further propose V-Score, a diagnostic question-answering metric that evaluates role-aware figurative faithfulness in generated images. Evaluations on recent high-performing TTI models show that vehicle intrusion persists across languages and figurative categories. As a lightweight mitigation, we introduce VISTA-Guard, which partially reduces vehicle intrusion and suggests a practical path toward more figuratively faithful TTI generation. All resources will be released publicly.

补充信息

↑