人类的第六感:多模态模型中直觉视觉推理的基准测试
Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models
浏览论文内容
中文总结 AI 辅助
提出HSS基准,测试多模态模型在图像和视频中的直觉视觉推理能力,发现前沿模型远逊于人类(93.1% vs 53.6%),即使智能体操作也未能完全弥合差距。
中文摘要 AI 辅助
人类在场景中感知到的信息远多于明确描绘的内容:一瞥即可捕捉过去的原因和未来的轨迹;快速一瞥即可判断车辆能否在两辆停放的汽车之间通过;几秒钟的视频即可揭示房间中谁拥有权威;而短暂的片段则能凸显诸如不成文规则或隐藏标签等微妙的抽象模式。这种能力体现了人类第六感的一种形式:一种超越原始感官知觉、恢复隐含信息的直觉推理机制。至关重要的是,这种快速、零样本的视觉直觉支撑着日常导航和社交互动,使其成为与人类一同部署的多模态大语言模型(MLLMs)的关键能力。然而,现有的视觉基准要么针对学术和数学领域的深思熟虑的专家级分析,要么针对低级感知,使得人类所进行的直觉推理在很大程度上未得到测试。为弥补这一空白,我们引入了“人类的第六感”(HSS),一个用于直觉视觉推理的基准。HSS涵盖多样的图像和视频输入,在结构化分类法下组织项目,并为每个项目配以人类编写的提示,探询人们一眼就能推断出的隐含时间、空间、社交和抽象结构。前沿MLLMs未达到人类表现:参与者达到93.1%的准确率,而最强的模型GPT-6-astra即使在最大推理努力下也仅达到53.6%。尽管在需要高级感知和知识的许多复杂任务中表现出色,当前模型在这些对人类而言直觉性的视觉任务上仍显著挣扎。我们进一步探索了将动态视觉操作应用于HSS的智能体设置,这缩小了差距但未完全消除。HSS将直觉视觉推理确立为一个可测量的轴,并引导关注当前基准扩展尚未触及的能力。
英文摘要
Humans perceive far more in a scene than what is explicitly depicted: a single glance captures past causes and future trajectories; a quick peek determines if a vehicle can fit between two parked cars; a few seconds of video reveals who holds authority in a room; and a fleeting clip highlights subtle abstract patterns like unwritten rules or hidden labels. This capacity reflects a form of humanity's sixth sense: an intuitive reasoning mechanism that recovers implicit information beyond raw sensory perception. Crucially, this rapid, zero-shot visual intuition underpins everyday navigation and social interaction, making it a vital capability for Multimodal Large Language Models (MLLMs) deployed alongside people. Existing visual benchmarks, however, target either deliberate expert-level analysis in academic and mathematical domains or low-level perception, leaving the intuitive reasoning that people perform largely untested. To bridge this gap, we introduce Humanity's Sixth Sense (HSS), a benchmark for intuitive visual reasoning. HSS spans diverse image and video inputs, organizes items under a structured taxonomy, and pairs each with human-written prompts probing the implicit temporal, spatial, social, and abstract structure that people infer at a glance. Frontier MLLMs fall short of human performance: participants reach 93.1% accuracy, while the strongest model, GPT-6-astra, reaches only 53.6% even at maximum reasoning effort. Despite excelling in many complex tasks that require advanced perception and knowledge, current models still struggle significantly on these visual tasks that are intuitive for humans. We further explore agentic setup that apply dynamic visual manipulation to HSS, which narrows but does not close the gap. HSS establishes intuitive visual reasoning as a measurable axis and directs attention to a capability that scaling on current benchmarks has so far left behind.
发表机构
- ScaleAI
- Elorian
- University of California, Santa Cruz(加州大学圣克鲁兹分校)
机构由 AI 辅助整理,请以论文原文为准。