arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhysMent:一种用于物理问题中LLM推理的交互式方法

PhysMent: An Interactive Approach For LLM Reasoning In Physics Problems

Joseph Chan, Utkarsh Jha, Xiyin Yang, Abhinav Jarajapu, Anik Sahai, Eddie Hu, Robin Jeshua Deepak, Stefano Saravalle, Aditya Shah

arXiv 2609.13152首次发表:更新:

AI 中文总结

PhysMent通过MuJoCo模拟器交互评估LLM物理推理,发现模型在定性任务上表现良好(高达80%),但在定量多步骤任务上准确率低于30%,瓶颈在于程序性工具使用而非概念理解。

AI 中文摘要

大型语言模型(LLM)在静态科学基准测试中表现强劲,但它们通过主动实验对物理世界进行推理的能力仍鲜为人知。我们引入了PhysMent,一个通过迭代式、工具介导的与MuJoCo物理模拟器交互来评估LLM物理推理的基准。与预先提供所有数量的静态基准不同,PhysMent要求模型在回答前通过施加力、查询物体状态、推进时间以及修改场景几何来发现信息。该基准包含105个经典力学场景,按四个难度级别(简单/困难及单概念/多概念)组织,涵盖三种场景模态(标准、物体创建、隐藏物体)以及一个场景操作类别,并使用六维评分框架进行评估。结果表明,当前模型在定性单概念任务上表现尚可(准确率高达80%),但在需要精确、多步骤实验程序的定量任务上大幅下降:大多数模型在最难的单概念类别中准确率低于30%,其瓶颈在于程序性(自适应多步骤工具使用)而非概念性负担。在七个模型中,准确率介于25%至67%之间,失败原因在于过早提交答案、探索效率低下以及未能一致地基于模拟器反馈进行推理,而非概念性缺陷。

英文摘要

Large language models (LLMs) perform strongly on static science benchmarks, yet their ability to reason about the physical world through active experimentation remains poorly understood. We introduce PhysMent, a benchmark that evaluates LLM physical reasoning via iterative, toolmediated interaction with a MuJoCo physics simulator. Unlike static benchmarks that supply all quantities upfront, PhysMent requires models to discover information by applying forces, querying object states, advancing time, and modifying scene geometry before answering. The benchmark comprises 105 scenes of classical mechanics, organized across four difficulty regimes (Easy/Hard and Single/Multi), three scene modalities (standard, object creation, hidden objects), and a scene-manipulation category, evaluated with a six-dimensional scoring framework. Results show that current models perform reasonably well on qualitative single-concept tasks (up to 80% accuracy) but degrade substantially on quantitative tasks that demand precise, multi-step experimental procedures: most models fall below 30% on the hardest single-concept category, where the bottleneck is procedural (adaptive multi-step tool use) rather than conceptual load. Across the seven models, accuracy ranges from 25% to 67%, with failures due to premature answer submission, inefficient exploration, and inconsistent grounding in simulator feedback rather than conceptual gaps.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑