arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhysMind:从视频到可执行世界的无训练物理推理框架

PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning

Chen Yang, Shenxiang Zeng, Haoyang Zhao, Zhouyuan Xu, Youquan He, Haoyu Li, Mingyi Deng, Jiansheng Fan, Chen Wang

arXiv 2608.04575首次发表:更新:

AI 中文总结

PhysMind是无训练的智能体框架,为每个视频构建可执行世界,在CLEVRER、Physion++数据集及反事实问题上的物理推理准确率显著优于现有视觉语言模型。

AI 中文摘要

从视频中进行可靠的物理推理需要理解物体如何运动、相互作用以及对干预措施的反应。现有视觉语言模型(VLMs)往往难以解释这些动态过程,并可靠地推理未来和反事实结果。我们引入PhysMind,这是一个无训练的智能体框架,为每个视频构建一个可重复使用、与问题无关的可执行世界。PhysMind通过物体分割、网格重建和6D姿态跟踪恢复时间一致的动态场景,然后拟合解析连续时间动态和潜在物理参数,无需展开时间步长模拟器。给定问题后,它会检查、继续或编辑该世界,并从所得轨迹和交互中给出答案。与使用相同VLM的直接思维链(CoT)推理相比,PhysMind在CLEVRER上的准确率提高了38.23个百分点,在Physion++上提高了8.08个百分点。在反事实问题上,它超过了评估的最强VLM基线GPT-5.5,提高了19.25个百分点。

英文摘要

Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions. Existing vision-language models (VLMs) often struggle to interpret these dynamics and reason reliably about future and counterfactual outcomes. We introduce PhysMind, a training-free agentic framework that constructs one reusable, question-agnostic executable world per video. PhysMind recovers a temporally consistent dynamic scene through object segmentation, mesh reconstruction, and 6D pose tracking, then fits analytic continuous-time dynamics and latent physical parameters without unrolling a time-stepped simulator. Given a question, it inspects, continues, or edits the world and answers from the resulting trajectories and interactions. Relative to direct chain-of-thought (CoT) reasoning with the same VLM, PhysMind improves accuracy by 38.23 points on CLEVRER and 8.08 points on Physion++. On counterfactual questions, it exceeds the strongest evaluated VLM baseline, GPT-5.5, by 19.25 points.

Comments27 pages, 18 figures. Project page: https://physmind.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑