arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

哪一个是香蕉人?评估视觉-语言模型在多轮语用理解中的表现

Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation

Alvin Wei Ming Tan, Ben Prystawski, Veronica Boyce

arXiv 2608.29571首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过迭代参照游戏测试发现,视觉-语言模型虽能利用上下文解释指代表达,但难以构建有效解释所需的相关上下文,缺乏高效语言协作的核心技能。

AI 中文摘要

对上下文的灵活适应以及共享的语用直觉是人类流畅对话的基础。迭代参照游戏(玩家通过语言反复挑选新颖指代对象)是测试智能体在多轮语言环境中进行上下文敏感语用推理能力的典型案例。我们在迭代参照游戏中,针对人类和视觉-语言模型(vision-language models)识别描述意图含义的能力进行了测试,通过改变提供上下文的数量、顺序和相关性来控制变量。人类表现始终良好,而被评估的模型虽能利用先前上下文解释人类的指代表达,但难以构建有效解释这些表达所需的相关上下文。我们的结果表明,被评估的模型缺乏高效语言协作所需的核心技能。

英文摘要

Flexible adaptation to context and shared pragmatic intuitions contribute to smooth human conversation. Iterated reference games---in which players repeatedly pick out novel referents using language---present a test case for agents' ability to perform context-sensitive pragmatic reasoning in multi-turn linguistic environments. We tested humans and vision--language models on their ability to identify the intended meaning of descriptions produced in iterated reference games, varying the provided context in terms of amount, order, and relevance. While humans performed well consistently, the models we evaluated could make use of prior context to interpret humans' referring expressions, but they struggled to build up the relevant context to interpret those expressions effectively. Our results suggest that the models we evaluated lack core skills needed for efficient linguistic collaboration.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑