发表机构
Moonshot AI(登月人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
介绍用于评估多模态大语言模型原子视觉感知能力的PerceptionBench基准,通过自下而上方法构建错误分类法及相关问题,测试16个前沿模型,发现原子感知待解决,该基准为衡量MLLM视觉感知边界提供标准。
AI 中文摘要
我们介绍了感知基准(PerceptionBench),这是一个专门设计用于评估多模态大语言模型(MLLM)原子视觉感知能力的基准。现有基准通常无法分离感知:整体评估将感知错误与推理或领域知识失败混为一谈,而应用驱动的基准仅涵盖由启发式设计塑造的狭窄、碎片化领域。为解决这些限制,PerceptionBench采用自下而上的方法:通过诊断前沿MLLM在42个现有基准响应中的最早失败点,构建错误分类法,其感知分支定义了十种原子感知能力。在此分类法指导下,构建了3000个答案简短明确的验证问题,每个问题分离一种能力,难度源于感知而非推理或知识。对16个前沿MLLM的基准测试结果表明,原子感知在很大程度上仍未解决——没有模型达到60%的准确率,与感知相关的幻觉平均是最弱的能力,且相似的总体分数掩盖了能力概况的巨大差异。因此,PerceptionBench为测量和诊断MLLM的视觉感知边界提供了能力级标准。
英文摘要
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.