发表机构
Xi’an Jiaotong-Liverpool University(西交利物浦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出PhysElite基准,含11586道奥林匹克级物理题及配套资源,测试18款MLLM发现最强者准确率仅33.7%,还开展分步流程评估诊断推理失败环节。
AI 中文摘要
要理解(多模态)大型语言模型在物理问题上的表现,需要能反映专家级物理推理难度与广度的基准测试。现有物理基准存在两个关键局限:一是缺乏高难度数据集,二是未全面覆盖视觉形式、知识点及分步解题过程。这导致模型在现有数据集上的表现无法完全代表其解决复杂物理问题的能力。为解决这些问题,本文提出PhysElite,这是一个用于奥林匹克级物理推理的大规模双语多模态基准。PhysElite包含11586道奥林匹克级题目,每道题配有对应示意图、中英双语分步解题推导过程及最终答案。本文对18款开源及闭源多模态大型语言模型(MLLM)进行基准测试,发现最强模型的答案准确率仅达33.7%。本文还开展了分步流程评估,以诊断模型在推理链中的失败环节。相关数据集已发布于该httpsURL。
英文摘要
Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning. Existing physics benchmarks remain limited in the following two important ways: (1) short of high-difficulty datasets, and (2) lack of comprehensive coverage of visual forms, knowledge points, and step-by-step solution processes. As a result, model performance on current datasets may not be fully representative of their ability to solve complex physics problems. To address these issues, we present PhysElite, a large-scale bilingual multimodal benchmark for Olympiad-level physics reasoning. PhysElite contains 11,586 Olympiad-tier problems. For each problem, we provide corresponding visual diagrams, step-by-step bilingual Chinese-English solution derivations, and the final answer. We benchmark 18 open-source and closed-source MLLMs, and find that even the strongest model reaches only 33.7% answer accuracy. We additionally conduct step-level process evaluation to diagnose where models fail in the reasoning chain. Our datasets are released at https://huggingface.co/datasets/physelite/PhysElite.
CommentsAnnual Conference on Neural Information Processing Systems (NeurIPS) 2026