arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WorldSolver:LLM智能体能否通过求解器生成来模拟物理动力学?

WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?

Siru Jiang, Yongzhe Lyu, Shuo Lu, Yubin Wang, Yuxiang Zhang, Yue Liao, Bin Wang, Jian Liang, Tieniu Tan

arXiv 2610.08720首次发表:更新:

发表机构

NLPR & MAIS, CASIA; PKU(中国科学院自动化研究所模式识别国家重点实验室与多模态人工智能系统全国重点实验室; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

WorldSolver提出一个包含168个物理模拟任务的基准,评估LLM智能体生成求解器的能力,发现现有前沿智能体在视觉和物理正确性上表现不佳,最高得分仅48.7%。

AI 中文摘要

基于LLM的智能体正日益推动科学和工程问题求解的进步,其中物理模拟作为一个具有挑战性且实用的测试平台,用于再现复杂的物理现象,并在具身智能、游戏和电影中有着广泛应用。作为此类模拟的核心,求解器计算动态系统状态随时间的变化。构建此类求解器需要物理理解以识别合适的模型、数学推理以表述底层动力学,以及软件工程将其实现为可执行代码,然而LLM智能体的这一能力仍未得到充分探索。为此,我们引入了WorldSolver,一个包含168个模拟任务的基准,这些任务源自61篇经典计算机图形学论文中的物理现象,涵盖7个物理领域。每个任务包含一个代码脚手架,为场景提供固定的模拟环境,而求解器的实现则留给智能体完成。具体而言,我们沿三个维度进行评估:执行检查(用于成功执行)、视觉保真度(用于在渲染模拟中再现预期动态行为)以及物理合理性(用于对生成动力学进行基于物理的验证)。对前沿智能体的实验表明,生成可执行的求解器本身就很困难,而满足视觉和物理正确性则更加困难。GPT-5.6-Sol和Claude-Opus-5的表现优于其他被评估的智能体,但总体得分分别仅为48.7%和46.7%。WorldSolver是迈向智能体求解器生成的早期一步,我们希望它能推动朝着能够忠实模拟动态物理世界的智能体取得进展。代码可在以下网址获取:https URL。

英文摘要

LLM-based agents are increasingly advancing scientific and engineering problem solving, with physics simulation emerging as a challenging yet practical testbed for reproducing complex physical phenomena with application in embodied AI, games and films. As the workhorse of such simulation, a solver computes how the state of a dynamic system evolves over time. Building such solvers requires physical understanding to identify appropriate models, mathematical reasoning to formulate the underlying dynamics, and software engineering to implement them as executable code, yet this capability of LLM agents remains underexplored. To this end, we introduce WorldSolver, a benchmark of 168 simulation tasks derived from physical phenomena in 61 classic computer graphics papers, spanning 7 physical domains. Each task contains a code scaffold that provides a fixed simulation environment for the scene, with the solver implementation left for the agent to complete. Specifically, we evaluate them along three dimensions: Execution Checks for successful execution, Visual Fidelity for reproducing the intended dynamic behavior in the rendered simulation, and Physical Plausibility for physics-grounded verification of the generated dynamics. Experiments on frontier agents reveal that producing executable solvers is difficult itself, and satisfying visual and physical correctness is even harder. GPT-5.6-Sol and Claude-Opus-5 perform comparatively better than the other evaluated agents, yet achieve overall scores of only 48.7% and 46.7%, respectively. WorldSolver is an early step toward agentic solver generation, and we hope it helps drive progress toward agents that can faithfully simulate the dynamic physical world. Code is available at https://github.com/sirujiang/WorldSolver.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑