arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估人工智能在基础物理问题解决中的能力

Assessing AI in Introductory Physics Problem Solving

Amir Bralin, N. Sanjay Rebello

arXiv 2607.14303首次发表:更新:

发表机构

Texas Tech University; Purdue University(德克萨斯理工大学; 普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究评估OpenAI的o4-mini模型解决基础物理问题的能力,通过解决《费曼物理学讲义》章节末尾问题展开,发现模型总体准确率约90%,性能受问题模态和难度影响,表明先进大语言模型虽能解决不少基础物理问题,但表现不均衡。

AI 中文摘要

推理或推理扩展模型是新一代能够解决复杂问题的大语言模型。为研究其在物理问题解决方面的能力,我们评估了OpenAI的o4-mini模型解决《费曼物理学讲义》中传统章节末尾问题的表现,这些问题涵盖本科物理课程核心主题。分析了跨模态和问题难度的性能。该模型解决问题的总体准确率约为90%,但性能很大程度上取决于表示形式:纯文本问题准确率更高(96%),而需要文本和图像协同解释的问题准确率为79%。随着问题难度从低到高增加,准确率也显著下降。这些结果表明,当前最先进的大语言模型可以解决许多标准的基础物理问题,但其性能仍不均衡,受问题模态和难度限制。

英文摘要

Reasoning or inference-scaling models are the new generation of Large Language Models (LLMs) capable of complex problem solving. To investigate their problem-solving capability in physics, we evaluated model o4-mini by OpenAI on solving traditional, end-of-chapter problems from Halliday and Resnick's "Fundamentals of Physics," spanning core topics in the undergraduate physics curriculum. Performance was analyzed across modality and problem difficulty. The model solved the problems with overall accuracy of about 90%, but performance depended strongly on representation: accuracy was much higher on text-only problems (96%) than on problems requiring coordinated interpretation of text and images (79%). Accuracy also declined significantly as the problem difficulty increased from low to medium to high. These results show that state-of-the-art LLMs can solve much of the standard introductory physics problems, but that their performance remains uneven and constrained by problem modality and problem difficulty.

CommentsWorking paper, submitted as a placeholder

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑