arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

前沿模型在物理学上的表现如何?专家重新评分揭示评估缺陷与领先基准的接近饱和

How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks

Ali Ansari, Haoran Sun, Andy Zeyi Liu, Mark Jabbour, Yongshan Ding, Steven Girvin, Yu He, Sohrab Ismail-Beigi, Aleksander Kubica, Owen D. Miller, Corey O'Hern, Vidvuds Ozolins, David Poland, A. Douglas Stone, Frank C. van den Bosch, Logan Wright, Navid Akbari, Santanu Antu, Kangle Cai, Andrew Calabrese-Day, Mateo Cárdenes Wuttig, Meng Cheng, Barry T. Chiang, Ali Ghorashi, Shouzhen Gu, Haoyang Huang, Zhibo Kang, Lukas Kienesberger, Hantian Liu, Charles Lomba, Zhongling Lu, Wenchao Ma, Rohin E. McIntosh, Evan McKinney, Ivan Rojkov, Xulei Sun, Yarone Meir Tokayer, Naveen Balaji Umasankar, Mira Varma, Leda Wang, Qimin Wang, Tyler Wang, Haoyu Wei, Jinming Yang, Jinchen Zhao, Sherlock Tingrui Zhao, Qinyuan Zheng, Jay S. Zou, Lucas Baker, Arman Cohan, John Sous

arXiv 2609.13009首次发表:更新:

发表机构

Yale University; Jump Trading Group; University of Cambridge; University of Southern California(耶鲁大学; Jump Trading 集团; 剑桥大学; 南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过专家审计六个物理基准,发现前沿模型分数被低估,修正后GPT-5.6-Sol的mean@4大幅提升,表明现有基准接近饱和,需更严格的专家验证评估。

AI 中文摘要

在领先的物理基准测试上报告的低分数,包括人工智能分析智能指数(2026)中列出的那些,表明前沿语言模型在高级物理方面仍存在困难,这是对其科学推理和定量问题解决能力的严格考验。然而,这种印象并不总是与领域专家在使用这些模型于其工作中的经验相符。我们通过评估前沿模型在六个广泛使用的物理基准上的表现,并与专家一起进行审计,重新审视这些报告的结果,重点关注具有可验证最终答案的纯文本问题。对于物理学的每个子领域,具有相关专业知识的教师和研究生研究人员仔细审查问题陈述、参考答案和模型响应,以区分真正的模型错误与评分者错误、错误的参考答案以及模糊或规定不充分的问题。大多数最初被评估为不正确的审计案例反映了这些基准问题,而非模型在物理推理中的错误。然后,我们请专家通过纠正错误的参考答案以及修复或排除有缺陷的问题来解决这些基准问题。我们发现,GPT-5.6-Sol在HLE-Physics上的测量mean@4从47.3%上升到78.7%,在CMT-Benchmark上从61.0%上升到87.2%,而其修正后的pass@4在保留的54个CritPt挑战上达到94.4%。修正后的分数是在专家审查后保留的评估子集上计算的。在UGPhysics、PRISM-Physics和PHYBench的审计子集上的分数在修正后也大幅上升。这些发现表明,当前基准严重低估了前沿模型解决结构良好的物理问题的能力。在这些封闭式任务上的接近饱和凸显了对更严格、专家验证的评估的需求。

英文摘要

Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still struggle with advanced physics, a demanding test of their scientific reasoning and quantitative problem-solving abilities. Yet this impression does not always align with domain experts' experiences using these models in their work. We revisit these reported findings by evaluating frontier models on six widely used physics benchmarks and auditing them with experts, focusing on text-only problems with verifiable final answers. For each subfield of physics, faculty and graduate researchers with relevant expertise carefully review problem statements, reference solutions, and model responses to distinguish genuine model errors from grader errors, incorrect reference solutions, and ambiguous or underspecified questions. Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models' physics reasoning. We then ask experts to address these benchmarking issues by correcting erroneous reference solutions and repairing or excluding flawed questions. We find that GPT-5.6-Sol's measured mean@4 rises from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, while its corrected pass@4 reaches 94.4% on the 54 retained CritPt challenges. Corrected scores are computed on the retained evaluation subsets following expert review. Scores on the audited subsets of UGPhysics, PRISM-Physics, and PHYBench also rise substantially after correction. These findings suggest that current benchmarks substantially understate frontier models' ability to solve well-posed physics problems. Near-saturation on these closed-ended tasks highlights the need for more demanding, expert-validated evaluations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑