AI 中文总结
针对多模态模型工具使用监督错位问题,提出ToolVision方法,通过能力对齐的SFT与RL监督优化工具使用学习,在多个视觉基准上实现性能提升并将公开发布相关资源。
AI 中文摘要
借助图像进行思考可让多模态模型通过代码调用视觉工具,弥补感知能力的不足。然而,主流的SFT后接RL的范式在每个阶段都会产生不同的监督错位问题:SFT本应教授如何使用工具,但更强教师的轨迹可能因较小学生模型无法可靠复现或利用的感知能力而成功,导致学生模仿工具调用模式却未学会如何让工具发挥作用;RL本应教授何时使用工具,但仅基于结果的奖励会使有缺陷的工具执行成为负担,从而抑制工具使用,而对每一条正确使用工具的轨迹给予笼统奖励则会鼓励有效但低效的操作。为解决这两种错位问题,我们提出ToolVision:在SFT阶段,多智能体流水线探索候选轨迹,包含学生规模模型的委员会对逐步证据增益进行评分,以排序和修剪搜索分支,仅保留成功执行且答案正确的轨迹用于SFT;在RL阶段前,ToolVision将学习者使用工具与不使用工具的表现进行对比,仅在工具能提供明显益处的问题上,对成功的工具使用给予奖励。两种信号均自动从公开任务数据构建,无需额外标注工具使用或必要性的人工注释。ToolVision-8B在全部7个主要基准上均优于其基础模型,在全部3个高分辨率基准上均超越Thyme-7B、CodeVision-8B和CodeDance-7B,在V*和HRBench 8K上的性能优于Qwen3-VL-32B-Thinking,我们将公开发布数据集和源代码。
英文摘要
Thinking with images allows a multimodal model to compensate for limited perception by invoking visual tools through code. Yet the prevailing SFT-then-RL recipe creates a different supervision misalignment at each stage. SFT is expected to teach how to use tools, but trajectories from stronger teachers may succeed through perceptual capabilities that a smaller student cannot reliably reproduce or exploit, causing the student to imitate tool-call patterns without learning how to make them useful. RL is expected to teach when to use tools, but outcome-only rewards make fallible tool execution a liability and suppress tool use, whereas a blanket bonus for every correct tool-using trajectory encourages valid but ineffective operations. To address these two misalignments, we introduce ToolVision. During SFT, a multi-agent pipeline explores candidate trajectories, and a committee including student-scale models scores stepwise evidence gain to rank and prune the search branches. Only successfully executed trajectories with correct final answers are retained for SFT. Before RL, ToolVision compares the learner's performance with and without tools, then rewards successful tool use only on questions where tools provide a clear benefit. Both signals are constructed automatically from public task data without additional human annotations of tool use or necessity. ToolVision-8B improves over its base on all seven main benchmarks, surpasses Thyme-7B, CodeVision-8B, and CodeDance-7B on all three high-resolution benchmarks, and outperforms Qwen3-VL-32B-Thinking on V* and HRBench 8K. We will publicly release the datasets and source code.
Comments18 pages, 13 figures