arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2507.04766cs.LGcs.AIcs.CL

ABench-Physics:通过高难度和动态物理问题对LLMs中的物理推理进行基准测试

ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems

  • Zhejiang University(浙江大学)
  • Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

Yiming Zhang, Yingfan Ma, Yanmei Gu, Zhengkai Yang, Yihong Zhuang, Feng Wang, Zenan Huang, Yuanyuan Wang, Chao Huang, Bowen Song, Cheng Lin, Junbo Zhao

更新

AI总结:

本文提出ABench-Physics基准,包含静态和动态物理问题,评估LLMs物理推理与泛化能力,发现显著性能差距。

AI中文摘要:

大型语言模型(LLMs)在数学和编程等领域表现出色,但其在物理方面的能力仍未得到充分探索和理解。物理提出了独特的挑战,不仅需要精确计算,还需要深入的概念理解和物理建模技能。现有基准往往因难度有限、选择题形式和静态评估设置而不足,无法捕捉物理建模能力。在本文中,我们介绍了ABench-Physics,这是一个新颖的基准,旨在严格评估LLMs的物理推理和泛化能力。ABench-Physics包含两个部分:Phy_A,一个包含400个研究生或奥林匹克级别问题的静态集合;以及Phy_B,一个包含100个问题的动态子集,配备自动变体引擎,用于测试模型在不断变化条件下的鲁棒性。所有问题都需要精确的数值答案,并具有严格的格式和容差约束。我们对几种最先进的LLMs的评估揭示了显著的性能差距,突显了物理推理中持续存在的局限性,尤其是在动态变体的泛化方面。ABench-Physics为推进LLMs中的科学推理提供了一个具有挑战性和诊断性的框架。

英文摘要:

Large Language Models (LLMs) have shown impressive performance in domains such as mathematics and programming, yet their capabilities in physics remain underexplored and poorly understood. Physics poses unique challenges that demand not only precise computation but also deep conceptual understanding and physical modeling skills. Existing benchmarks often fall short due to limited difficulty, multiple-choice formats, and static evaluation settings that fail to capture physical modeling ability. In this paper, we introduce ABench-Physics, a novel benchmark designed to rigorously evaluate LLMs' physical reasoning and generalization capabilities. ABench-Physics consists of two components: Phy_A, a static set of 400 graduate- or Olympiad-level problems; and Phy_B, a dynamic subset of 100 problems equipped with an automatic variation engine to test model robustness across changing conditions. All questions require precise numerical answers, with strict formatting and tolerance constraints. Our evaluation of several state-of-the-art LLMs reveals substantial performance gaps, highlighting persistent limitations in physical reasoning, especially in generalization to dynamic variants. ABench-Physics provides a challenging and diagnostic framework for advancing scientific reasoning in LLMs.

↑