arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhysFieldBench:多模态模型能否理解物理场?

PhysFieldBench: Can Multimodal Models Understand Physical Fields?

Yuezhou Ma, Huikun Weng, Jialong Wu, Chenyi Zhao, Hang Zhou, Haonan Shangguan, Jianmin Wang, Mingsheng Long

arXiv 2609.34072首次发表:更新:

发表机构

Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出PhysFieldBench基准,评估多模态模型理解物理场的能力,发现零样本性能低,而监督视觉变换器表现更好,并探索后训练方法以提升泛化。

AI 中文摘要

多模态大语言模型(MLLMs)正日益被视为科学和工程智能体的核心组件,然而它们解释物理场的能力仍然鲜为人知。现有的物理基准大多侧重于教科书式的问题求解或直觉物理推理,而未解答MLLMs能否从连续场观测中推断出具有物理意义的信息。我们引入了PhysFieldBench,一个包含24个任务和1160个评估示例的基准,覆盖受控方程场、模拟物理场和观测物理场。这些任务评估三种推理形式:识别物理机制、比较潜在控制变量和预测结果属性。在具有代表性的开源和专有MLLMs中,零样本性能较低:最佳模型的机遇归一化得分为29.3,而几个开源模型仍接近随机水平。相比之下,一个任务特定的监督视觉变换器表现明显更好,表明输入包含可学习的物理信息。为了诊断这些失败,一项结构化的自我解释分析将大多数错误归因于遗漏的视觉模式和错误的视觉到物理映射。此外,为了探索后训练能否改善物理推理并泛化到未见任务,我们比较了监督微调(使用最终答案或思维链监督)和强化学习。最终答案监督整体表现最佳,但迁移效果较差,而思维链监督后的强化学习实现了最佳泛化。总之,这些发现强调了改进视觉到物理的接地和跨任务泛化,以使MLLMs在科学和工程工作流中可靠地解释物理场的必要性。

英文摘要

Multimodal large language models (MLLMs) are increasingly envisioned as core components of scientific and engineering agents, yet their ability to interpret physical fields remains poorly understood. Existing physics benchmarks largely emphasize textbook problem solving or intuitive physical reasoning, leaving open whether MLLMs can infer physically meaningful information from continuous field observations. We introduce PhysFieldBench, a benchmark comprising 24 tasks and 1,160 evaluation examples across controlled equation fields, simulated physical fields, and observed physical fields. The tasks assess three forms of inference: identifying physical mechanisms, comparing latent control variables, and predicting outcome properties. Across representative open-source and proprietary MLLMs, zero-shot performance is low: the best model achieves a chance-normalized score of 29.3, while several open-source models remain near chance. In contrast, a task-specific supervised vision transformer performs substantially better, demonstrating that the inputs contain learnable physical information. To diagnose these failures, a structured self-explanation analysis attributes most errors to missed visual patterns and incorrect visual-to-physical mappings. Further, to explore whether post-training can improve physical inference and generalize to unseen tasks, we compare supervised fine-tuning with final answers or chain-of-thought supervision and reinforcement learning. Final-answer supervision performs best overall but transfers less effectively, whereas reinforcement learning after chain-of-thought supervision achieves the best generalization. Together, these findings highlight the need to improve visual-to-physical grounding and cross-task generalization for MLLMs to reliably interpret physical fields in scientific and engineering workflows.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑