发表机构
Psychological & Brain Sciences, University of California, Santa Barbara; Department of Computer Science, University of California, Santa Barbara; Department of Electrical and Computer Engineering, University of California, Santa Barbara(加利福尼亚大学圣巴巴拉分校心理与脑科学系; 加利福尼亚大学圣巴巴拉分校计算机科学系; 加利福尼亚大学圣巴巴拉分校电气与计算机工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究十年间视觉语言模型进展,引入CSB数据集,评估模型在其上及MS-COCO样本中的准确性与视觉认知错误类型,发现MLLM消除简单与复杂场景描述准确性差距,几乎消除多数错误类型,为模型发展提供全面评估。
AI 中文摘要
在过去十年中,视觉语言模型(VLM)在视觉推理方面取得了显著进展。大多数评估使用的是简单场景(MS-COCO),未展示复杂的人类互动或行为,仅以少数未经策划的人类描述作为基准,且未关注模型的错误类型。本文引入了包含100张描绘复杂社会互动/行为图像的复杂社会行为(CSB)数据集。分析了2017年至2025年十年间VLM(四个预多模态大语言模型、MLLM和五个MLLM)场景描述的进展。在CSB数据集和MS-COCO样本上评估了模型和20个人类描述相对于黄金标准的准确性。分析了五种视觉认知错误类型:对象检测、识别、幻觉、场景理解和空间依赖性。CSB数据集在场景描述准确性方面比MS-COCO有更显著的提高,预MLLM的准确性远低于排名垫底的人类描述,而MLLM的准确性与排名靠前的人类描述相似。表明MLLM消除了简单MS-COCO场景和描绘复杂行为(CSB)场景之间在场景描述准确性上的差距。MLLM几乎消除了测试数据集中的所有错误类型,除了偶尔在场景描述中依赖与人类不同的图像区域(空间依赖性错误)。还表明检测、识别和幻觉错误对场景描述准确性影响最大。这些发现更全面地评估了视觉语言模型在过去十年中的进展。
英文摘要
Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handful of non-curated human descriptions as a benchmark, and have not focused on understanding the model's error types. Here, we introduce the Complex Social Behavior (CSB) dataset, containing 100 images depicting complex social interactions/behaviors. We analyze the progression of scene descriptions over a decade (2017-2025) of VLMs (four pre-Multimodal Large Language Models, MLLMs, and five MLLMs). We evaluate the accuracy of the models and 20 human descriptions relative to a gold standard on the CSB dataset and on a sample from MS-COCO. We analyzed five visual-cognitive error types: object detection, recognition, hallucination, scene understanding, and spatial dependence. The CSB dataset showed a more pronounced improvement than MS-COCO in scene description accuracy, with pre-MLLMs achieving much lower accuracy than the bottom-ranked human descriptions and MLLMs attaining accuracies similar to the top-ranked human descriptions. We show that MLLMs have eliminated the gap in scene description accuracy between simpler MS-COCO scenes and scenes depicting complex behaviors (CSB). MLLMs have almost eliminated all error types in our tested datasets, except for occasionally relying on different image regions for scene descriptions than humans do (spatial dependence error). We also show that detection, recognition, and hallucination errors have the highest impact on scene description accuracy. Together, our findings provide a more thorough evaluation of how visual language models have advanced over the last decade.