发表机构
East China Normal University(华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人员推出OmniPhys这一覆盖中文教育语料库中中学至大学物理问题的大规模多模态基准,含15246个问题与19850张图像,可评估MLLMs的物理理解、推理及结构化图表生成能力,发现当前模型存在复杂推理与视觉生成差距,将推进物理领域多模态智能发展。
AI 中文摘要
多模态大语言模型(MLLMs)已展现出解决各类视觉与文本推理任务的强大能力,但其在物理领域的发展因缺乏综合性基准而受到显著阻碍。为填补这一空白,我们推出OmniPhys——一个用于多模态物理理解与推理的大规模基准,覆盖中文教育语料库中从中学到大学级别的问题。OmniPhys包含15246个问题和19850张图像,附带详细注释,支持对推理过程和知识使用进行细粒度分析。除常规评估外,OmniPhys是一个系统评估物理领域多模态输出的基准,包括模型生成结构化物理图表的能力,而结构化物理图表是真实物理问题解决的基础组成部分。大量评估显示,当前MLLMs的能力存在关键差距,尤其在复杂推理和视觉生成方面。为解决这一问题,我们发布OmniPhys,作为推进物理及科学领域多模态智能的基础资源,代码与数据可在该https URL获取。
英文摘要
Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks. However, their development in the physics domain is significantly hindered by the lack of a comprehensive benchmark. To fill this gap, we introduce OmniPhys, a large-scale benchmark for multimodal physics understanding and reasoning, covering middle school through university-level problems from Chinese Educational Corpora. OmniPhys consists of 15,246 questions and 19,850 images, accompanied by detailed annotations that support fine-grained analysis of reasoning processes and knowledge usage. Beyond conventional evaluation, OmniPhys is a benchmark that systematically evaluates multimodal outputs in the physics domain, including models' ability to generate structured physics diagrams, which constitute a fundamental component of authentic physics problem solving. Extensive evaluations reveal critical gaps in the capabilities of current MLLMs, especially in complex reasoning and visual generation. To address this, we release OmniPhys to serve as a foundational resource for advancing multimodal intelligence in physics and scientific domains. Codes and data are available at https://github.com/ECNU-RAIL/OmniPhys-EMNLP2026.
CommentsAccepted to Findings of EMNLP 2026