arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

感知基准:评估多模态大语言模型中的原子视觉感知

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang, Jia Li, Haoning Wu, Yuhao Dong, Zuhao Yang, Jinguo Zhu, Haoyu Lu, Zijia Zhao, Tongtian Yue, Zhangyang Qi, Junwei Yang, Mengfan Dong, Peizhou Cao, Chenzhuang Du, Zaida Zhou, Haotian Yao, Hao Yang, Hongcheng Gao, Lin Sui, Weihong Li, Xinxing Zu, Jia Chen, Yao Wang, Xiaoxue Wu, Yalin Wang, Y. Charles, Yiping Bao, Yangyang Liu, Zhiqi Huang, Xinyu Zhou

arXiv 2607.24957首次发表:更新:

发表机构

Moonshot AI(登月人工智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

介绍用于评估多模态大语言模型原子视觉感知能力的PerceptionBench基准,通过自下而上方法构建错误分类法及相关问题,测试16个前沿模型,发现原子感知待解决,该基准为衡量MLLM视觉感知边界提供标准。

AI 中文摘要

我们介绍了感知基准(PerceptionBench),这是一个专门设计用于评估多模态大语言模型(MLLM)原子视觉感知能力的基准。现有基准通常无法分离感知:整体评估将感知错误与推理或领域知识失败混为一谈,而应用驱动的基准仅涵盖由启发式设计塑造的狭窄、碎片化领域。为解决这些限制,PerceptionBench采用自下而上的方法:通过诊断前沿MLLM在42个现有基准响应中的最早失败点,构建错误分类法,其感知分支定义了十种原子感知能力。在此分类法指导下,构建了3000个答案简短明确的验证问题,每个问题分离一种能力,难度源于感知而非推理或知识。对16个前沿MLLM的基准测试结果表明,原子感知在很大程度上仍未解决——没有模型达到60%的准确率,与感知相关的幻觉平均是最弱的能力,且相似的总体分数掩盖了能力概况的巨大差异。因此,PerceptionBench为测量和诊断MLLM的视觉感知边界提供了能力级标准。

英文摘要

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑