主动观察者的测试
An Exam for Active Observers
- University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究多模态大语言模型是否具备主动观察能力,引入ActiveVision基准测试,发现前沿模型表现不佳,即便能编写视觉代码差距仍存,表明当前模型缺乏主动视觉观察,呼吁构建闭合感知 - 推理循环的架构和训练目标。
AI中文摘要:
人类视觉是一个闭环:注视会不断被中间假设重新引导,而非单一快照。数十年来,心理物理学和认知科学认为这种主动观察对广泛任务至关重要。当前视觉语言基准无法回答当今的多模态大语言模型(MLLM)是否进行主动观察。我们引入了ActiveVision基准,它能让MLLM的主动观察可测量,包含3类17个任务,旨在促使重复视觉感知。前沿MLLM在ActiveVision上表现不佳,最高得分模型GPT - 5.5在最高推理努力层级仅解决10.6%的项目,17个任务中有11个得零分,Claude Fable 5也仅解决3.5%,远落后于平均96.1%的三名人类参与者。即便模型编写并运行自己的视觉代码,差距依然存在,因为代码在现实图像中不可靠,而捕捉其失败需要模型所缺乏的主动感知。这些结果表明当前MLLM缺乏强大的主动视觉观察能力,这促使构建能闭合感知 - 推理循环的架构和训练目标。
英文摘要:
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception rather than a single static description. Frontier MLLMs collapse on ActiveVision: the highest-scoring model we evaluate, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores zero on 11 of the 17 tasks, and even Claude Fable 5, despite topping most reasoning and coding leaderboards, solves just 3.5%, far behind three human participants who average 96.1%. Furthermore, much of the gap persists even when models write and run their own vision code: such code is unreliable on realistic imagery, and catching its failures itself requires the active perception the models lack. Together, these results indicate that current MLLMs lack robust active visual observation, motivating architectures and training objectives that close the perception-reasoning loop.