歧义去了哪里?探究多模态模型如何解释多义词
Where did the ambiguity go? Examining how multimodal models interpret polysemous words
浏览论文内容
中文总结 AI 辅助
该研究对比17个文本到图像模型和15个文本生成模型,发现多模态模型生成图像的词义多样性低于文本,揭示了基础模型在不同模态间意义表达的迁移 gap。
中文摘要 AI 辅助
人类语言具有高度多义性,许多常见词汇(如“bank”或“palm”)承载着不同含义,这些含义影响人类的交流与想象。已有研究表明大型语言模型(LLMs)能理解这种语义多样性,但人们对多义性在图像等其他模态中的表现知之甚少。本研究针对17个文本到图像模型和15个文本生成模型,为每个模型提供无上下文限定含义的多义词,通过大量采样测量模型生成的词义。研究发现存在明显的多模态 gap:在每个模型家族中,生成图像的词义数量远少于生成句子(归一化熵分别为0.10和0.25),且两者的多样性都远低于人类对相同词汇的想象(归一化熵为0.47)。然而,当要求模型列出对应词汇各可能含义的输出频率时,其预测的分布比实际输出空间更多样。这些结果揭示了基础模型表达意义时存在多模态 gap,且其理解无法在不同模态间忠实且平等地迁移。
英文摘要
Human language is highly polysemous. Many common words (e.g., "bank" or "palm") carry several distinct meanings that shape what humans communicate and imagine. Large language models (LLMs) have been shown to understand this multiplicity of meaning, but much less is known about how polysemy surfaces in other modalities such as images. We study this across 17 text-to-image and 15 text-generation models by giving each a polysemous word with no context to fix its meaning and measuring which senses are produced over many samples. We find a clear multimodal gap, where within every model family, generated images settle on far fewer senses than generated sentences (normalized entropy 0.10 vs. 0.25), and both are far less varied than what people imagine for the same words (normalized entropy 0.47). However, when we instead ask a model to list how often it would generate outputs corresponding to each possible meaning of a word, it predicts distributions that are more diverse than the actual space of outputs. These results reveal a multimodal gap in how foundation models express meaning, and how their understanding may not transfer faithfully nor equally across modalities.
发表机构
- Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。