arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

新证据,相同选择:测试视觉语言模型中的物理实验选择

New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models

Sourajit Saha, Shubhashis Roy Dipta, Nobin Sarwar, Shaswati Saha, Yuxuan Jiang, Siyuan Li, Qiheng Wang

arXiv 2609.11022首次发表:更新:

发表机构

University of Maryland, Baltimore County; University of Georgia(马里兰大学巴尔的摩县分校; 佐治亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过受控实验基准测试视觉语言模型在物理推理中能否根据新证据选择最优实验或直接回答,发现多数模型在行动切换上表现不佳,仅少数能正确决策。

AI 中文摘要

模型首先看到来自一个物理测量实验的图像,例如一个木块滑行了多远,然后必须回答关于一次新试验的问题,例如在固定推力后木块是否会通过一个目标。初始实验可能提供足够的信息来回答问题,或者模型可能需要另一个测量,例如物体的质量、摩擦力、恢复系数或弹簧刚度。我们研究视觉语言模型能否决定何时立即回答,以及当需要更多证据时,选择执行哪个实验。当前的物理推理基准通常只评估最终答案,因此它们不能直接衡量这种决策能力。我们引入了一个受控评估,其中每个问题提供一个测量图像和四个可能的物理世界,这些世界由两个可能的质量和另一个相关属性的两个可能值组合而成。模型必须要么停止并回答,要么选择能够解决该问题的最便宜的额外实验。我们构建了匹配的问题对,其中改变观察到的测量或问题会改变最优行动。由于所有可能的世界和实验成本都是已知的,我们可以明确确定最优选择。在六个开放模型和144个物理参数集上,直接响应在95.1%到100%的图像对中重复相同的行动,即使正确的行动发生了变化。简短的推理改善了行动切换,但最好的模型仅在5.9%的图像对上正确做出了两个决策。额外的分析揭示了在测量解释、物理推理和响应格式方面的失败。通过将证据选择与最终答案分开评估,我们的基准揭示了传统答案准确性可能忽视的物理推理局限性。

英文摘要

A model first sees an image from one physical measurement experiment, such as how far a block coasted, and must answer a question about a new trial, such as whether the block will pass a target after a fixed push. The initial experiment may provide enough information to answer, or the model may need another measurement, such as the object's mass, friction, restitution, or spring stiffness. We study whether vision language models can decide when to answer immediately and, when more evidence is needed, which experiment to perform. Current physical reasoning benchmarks usually evaluate only the final answer, so they do not directly measure this decision-making ability. We introduce a controlled evaluation where each problem provides one measurement image and four possible physical worlds created by combining two possible masses and two possible values of another relevant property. The model must either stop and answer or select the cheapest additional experiment that can resolve the question. We construct matched problem pairs where changing either the observed measurement or the question changes the optimal action. Since all possible worlds and experiment costs are known, we can explicitly determine the optimal choice. Across six open models and 144 physical parameter sets, direct responses repeat the same action for 95.1% to 100% of image pairs even when the correct action changes. Brief reasoning improves action switching, but the best model makes both decisions correctly for only 5.9% of image pairs. Additional analysis reveals failures in measurement interpretation, physical reasoning, and response formatting. By evaluating evidence selection separately from final answers, our benchmark reveals limitations in physical reasoning that conventional answer accuracy can overlook.

CommentsUnder Review at PhysWorldAI @ NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑