发表机构
Shanghai Artificial Intelligence Laboratory; Shanghai Jiao Tong University; Shanghai Innovation Institute(上海人工智能实验室; 上海交通大学; 上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过因果审查方法,发现多模态大语言模型的视觉工具使用存在“调用而不查看”“查看而不规划”等错觉,整体准确率增益下大量场景无因果效果。
AI 中文摘要
“图像思考”范式为多模态大语言模型(multimodal LLMs)配备了裁剪缩放等主动视觉操作,但使用这些操作的模型在 token 成本显著更高的情况下,仅能获得边际收益甚至负收益,还可能反复裁剪无关区域,在直接推理能正确回答的问题上失败。本文研究返回的视觉证据是否会对答案产生因果影响,为此将视觉工具使用构建为因果图,区分观测介导路径与动作诱导捷径,并从三个层面进行干预审查:策略层面(对比工具使用与直接推理)、轨迹层面(在展开过程中破坏所有观测)、步骤层面(在固定前缀下反事实替换单个观测)。本文提出的步骤层面估计量为视觉证据增益,用于分离每个返回观测的贡献。在六个代表性模型和五个细粒度感知基准上,研究发现策略校准存在两种失败模式:“调用而不查看”即返回观测对答案无因果影响,“查看而不规划”即观测有信息但调用顺序不连贯。轨迹层面诊断分解了策略层面的准确率增益,显示增益集中在少数校准良好的情况中,本文将这种差异称为视觉工具使用的错觉:尽管整体准确率有增益,但视觉工具使用在大量展开过程中并无因果效果。代码可在该 https URL 获取。
英文摘要
The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on questions that direct inference answers correctly. We ask whether the returned visual evidence causally affects the answer. To answer this question, we formulate visual tool-use as a causal graph that separates observation-mediated paths from action-induced shortcuts. We then audit it through interventions at the three levels: policy (comparing tool-use with direct inference), trajectory (corrupting all observations during rollout), and step (counterfactually replacing one individual observation under a fixed prefix). Our step-level estimand, Visual Evidence Gain, isolates the contribution of each returned observation. Across six representative models and five fine-grained perception benchmarks, we uncover policy miscalibration with two failure modes. In Calling Without Looking, returned observations have no causal effect on the answer. In Looking Without Planning, observations are informative but the call schedule is incoherent. A trajectory-level diagnostic decomposes the policy-level accuracy gain into per-group contributions and shows that the gain is concentrated in a Calibrated minority. We term this discrepancy the illusion of visual tool-use: despite aggregate accuracy gains, visual tool-use is not causally effective across a broad range of rollouts. The code is available at https://github.com/OpenCausaLab/CauAudit.
CommentsEMNLP 2026 Findings