存在并非忠实:文生图中的喻体侵入
Presence Is Not Faithfulness: Figurative Vehicle Intrusion in Text-to-Image Generation
浏览论文内容
中文总结 AI 辅助
针对文生图模型将喻体误渲染为可见物体的问题,提出VISTA基准和V-Score指标,并引入VISTA-Guard缓解策略,以提升比喻忠实度。
中文摘要 AI 辅助
文生图(TTI)模型日益能够根据自然语言提示生成高质量图像,然而比喻性语言暴露了一个失败:本应引导本体描绘的喻体,反而可能被渲染为可见物体。我们将此失败称为“喻体侵入”:侵入内容在文本上得到许可,但被赋予了错误的视觉角色,这表明视觉存在并不总是忠实,且以存在为导向的评估可能遗漏此类错误。为系统研究此现象,我们引入了“喻体侵入与语义本体评估”(VISTA),一个按比喻形式和映射机制组织的多语言比喻提示基准。我们进一步提出V-Score,一种诊断性问答指标,用于评估生成图像中角色感知的比喻忠实度。对近期高性能TTI模型的评估显示,喻体侵入在多种语言和比喻类别中持续存在。作为一种轻量级缓解措施,我们引入VISTA-Guard,它能部分减少喻体侵入,并为更比喻忠实的TTI生成提供了一条实用路径。所有资源将公开发布。
英文摘要
Text-to-image (TTI) models increasingly generate high-quality images from natural-language prompts, yet figurative language exposes a failure: a vehicle that should guide the depiction of a tenor may instead be rendered as a visible object. We call this failure Figurative Vehicle Intrusion: the intruding content is textually licensed, but it is assigned the wrong visual role, showing that visual presence is not always faithfulness and that presence-oriented evaluation can miss such errors. To study it systematically, we introduce Vehicle Intrusion and Semantic Tenor Assessment (VISTA), a multilingual benchmark of figurative prompts organized by Figurative Form and Mapping Mechanism. We further propose V-Score, a diagnostic question-answering metric that evaluates role-aware figurative faithfulness in generated images. Evaluations on recent high-performing TTI models show that vehicle intrusion persists across languages and figurative categories. As a lightweight mitigation, we introduce VISTA-Guard, which partially reduces vehicle intrusion and suggests a practical path toward more figuratively faithful TTI generation. All resources will be released publicly.