arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12575cs.CLcs.AIcs.HC

多模态语言模型中的校准模糊性:人类触及文化引用,而模型描述图片

Calibrated Ambiguity in Multimodal Language Models: Humans reach for cultural references, while models describe the picture

  • The Alan Turing Institute(艾伦·图灵研究所)
  • Technical University of Darmstadt(达姆施塔特工业大学)
  • University of Southampton(南安普顿大学)
  • University of Chicago(芝加哥大学)
  • University of Sheffield(谢菲尔德大学)
  • University of Edinburgh(爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

Cody Kommers, Mingrui Ye, Evelyn Gius, Daniela Mihai, Hoyt Long, Zheng Yuan, Drew Hemment

AI总结:

本研究通过桌游《画物语》任务,比较人类与多模态语言模型生成的线索,发现模型存在模糊性坍缩和文化扁平化问题,即输出过度具体且缺乏文化引用。

AI中文摘要:

模糊性常被视为人工智能系统需要解决的缺陷——但在人类交流与文化中,模糊性也可以是一种生成性资源。从幽默到政治再到艺术,人们用开放到足以引发多种解读、又受约束到可被理解的文字和图像来表达自己。我们通过源自桌游《画物语》的任务,将这种校准模糊性的概念操作化。基于一套新颖的校准模糊性编码规则,我们比较了人类与多模态语言模型生成的线索差异,发现模型始终表现出模糊性坍缩(即其输出过度具体化,没有为多种合理解读留下空间)。与人类线索不同,AI生成的线索还表现出文化扁平化;即使被提示使用典故和比喻性语言,它们也几乎从不引用文化情境知识。

英文摘要:

Ambiguity is often treated as a bug for AI systems to resolve---but in human communication and culture, ambiguity can also be a generative resource. From humour to politics to art, people express themselves in words and images that are open enough to invite different interpretations, yet constrained enough to be interpretable. We operationalise this notion of calibrated ambiguity with a task drawn from the parlour game Dixit. We compare differences in clues generated by human vs multimodal language models, based on a novel coding rubric for calibrated ambiguity, and find that models consistently exhibit ambiguity collapse (i.e., their outputs are over-specified, leaving no room for multiple legitimate interpretations). Unlike human clues, AI-generated clues also exhibit cultural flattening; they almost never make reference to culturally-situated knowledge, even when prompted to use allusion and figurative language.

补充信息

↑