语言模型无需视觉输入即可进行想象?Ekphrasis:测量仅文本大型语言模型的视觉创意构想能力
Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs
浏览论文内容
中文总结 AI 辅助
本研究推出基准Ekphrasis,测量仅文本大型语言模型的视觉创意构想能力,发现该能力与流畅度分离,且其排序可跨模态稳定存在。
中文摘要 AI 辅助
当前评估无法分离仅文本的语言模型是否能在图像生成前自主产生视觉概念。流畅的视觉散文可能掩盖视觉规划的缺陷:一个答案看似有创意,实则重复熟悉的视觉陈词滥调或未能指定可渲染的场景。我们将视觉创意构想(Visual Creative Ideation, VCI)定义为生成有用、富有表现力且具有群体新颖性的文本视觉规划的能力,并推出Ekphrasis,这是一个包含400项任务的基准,涵盖抽象、组合、转换和适应四个类别。Ekphrasis通过特定维度的检查表对匿名成对比较进行评分,使用Bradley-Terry模型汇总偏好,并利用类型化概念图将特定任务的群体陈词滥调转换为新颖性参考。在14种语言模型中,VCI将有用性、表现力和新颖性区分开来,而非简化为流畅度:强大的模型通过不同的表现轮廓获得相似的整体分数,且有用的规划仍可能是视觉陈词滥调。一项跨模态 grounding 研究进一步表明,文本层面的VCI排序在忠实渲染和盲图像层面偏好判断中基本保持一致,支持Ekphrasis作为超越散文质量的视觉构想测量工具。
英文摘要
Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar visual clichés or failing to specify a renderable scene. We define Visual Creative Ideation (VCI) as the ability to produce textual visual plans that are useful, expressive, and population-novel, and introduce Ekphrasis, a 400-task benchmark spanning Abstraction, Combination, Transformation, and Adaptation. Ekphrasis scores anonymized pairwise comparisons with dimension-specific checklists, aggregates preferences with Bradley-Terry models, and uses Typed Idea Graphs to convert task-specific population clichés into novelty references. Across 14 language models, VCI separates usefulness, expressiveness, and novelty rather than reducing to fluency: strong models achieve similar overall scores through different profiles, and useful plans can remain visually clichéd. A cross-modal grounding study further shows that text-level VCI ordering largely survives faithful rendering and blind image-level preference judgment, supporting Ekphrasis as a measure of visual ideation beyond prose quality.
发表机构
- The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。