AI 中文总结
本文提出See2Think评估框架,通过See2ThinkBench和VAoT测试多模态模型对中间视觉状态的依赖,发现视觉推理具模型与环境依赖性,准确渲染是瓶颈,受损反馈会使模型准确率降超10个百分点。
AI 中文摘要
多模态大语言模型在推理过程中越来越多地使用草图、注释、工具和中间图像,但目前尚不清楚它们是否真正依赖这些视觉状态。现有基准存在两方面局限:一是任务集合覆盖范围狭窄,或存在部分可仅通过文本解决的样本;二是评估侧重最终答案,未诊断中间视觉状态的生成、渲染及使用方式。本文提出See2Think,这是一个统一评估框架,包含See2ThinkBench和视觉思维链(Visual Action-of-Thought, VAoT)。See2ThinkBench涵盖12类任务共1200个开放式、视觉依赖问题,涉及2D结构化、3D场景及现实世界推理。VAoT在四种受控推理设置下记录文本思维、视觉操作、渲染状态及后续推理。通过评估代表性专有与开源多模态模型,发现视觉推理具有强烈的模型和环境依赖性,无单一设置在所有任务中持续占优。过程分析进一步显示,模型通常会选择相关视觉操作,但准确渲染仍是最明显的瓶颈,高反馈获取未必转化为准确率提升。在任务相关的受损反馈下,模型表现出对视觉状态的行为依赖,受控干预下准确率下降超过10个百分点。
英文摘要
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over 10 percentage points in controlled interventions.
Comments10 pages, 5 figures, and 8 tables