发表机构
SenseTime Research(商汤科技研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出OpenVisTool框架,构建含因果监督的视觉工具使用轨迹数据集OpenVisTool-42K,微调后模型视觉工具使用性能提升,大模型接近领先闭源系统。
AI 中文摘要
视觉工具使用已成为多模态智能体主动获取固定图像编码之外证据的一项基本能力。主流方案从教师生成的轨迹中学习该能力,这些轨迹会按答案正确性进行筛选,隐含假设是每一个成功的演示都能提供有效监督。我们认为该假设存在缺陷:强大的教师往往无需调用工具就能得出正确答案,模仿此类轨迹会让学生认为工具调用伴随正确答案,而非工具观察为答案提供依据。我们提出OpenVisTool,这是一个用于构建具有指导性的视觉工具使用轨迹的开源框架,可为工具学习提供有效监督。核心见解在于,仅当轨迹满足答案正确(结果有效性)且工具观察对该答案具有因果贡献(因果效用)时,才应保留该轨迹。该框架分为三个阶段:难度筛选,选择不借助工具就无法可靠回答的查询;特定领域的轨迹合成,引出连贯的工具使用轨迹;监督验证,联合测试上述两个条件。该框架不鼓励模型模仿工具调用,而是让模型学习何时以及如何获取视觉证据。我们利用该框架构建了涵盖五个视觉推理领域的数据集OpenVisTool-42K,以及覆盖相同领域的基准OpenVisTool-Bench。在四个主干模型(4B-27B)上,对OpenVisTool-42K进行微调后,视觉工具使用性能持续提升,且在两个分布外基准上也取得了增益;规模较大的模型已接近领先的闭源系统。上述证据表明,有效的视觉工具使用是通过因果基础的监督而非工具调用模式习得的。
英文摘要
Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for answer correctness, implicitly assuming that every successful demonstration provides effective supervision. We argue this assumption is flawed: a strong teacher often reaches the correct answer without needing its tool calls, and imitating such trajectories teaches a student that tool calls accompany correct answers, not that tool observations ground them. We present OpenVisTool, an open framework for constructing instructive visual tool-use trajectories that provide effective supervision for tool learning. The key insight is that a trajectory should be retained only if its answer is correct (outcome validity) and its tool observations causally contribute to that answer (causal utility). The framework operates in three stages: difficulty screening to select queries that are not reliably answerable without tools, domain-specific trajectory synthesis to elicit coherent tool-use trajectories, and supervision verification to jointly test both conditions. Rather than encouraging models to imitate tool calls, the resulting supervision teaches when and how visual evidence should be acquired. Using this framework, we construct OpenVisTool-42K, a dataset spanning five visual reasoning domains, together with OpenVisTool-Bench, a benchmark covering the same domains. Across four backbones (4B-27B), fine-tuning on OpenVisTool-42K consistently improves visual tool-use performance and yields gains on two out-of-distribution benchmarks; the larger models approach leading closed-source systems. The evidence suggests that effective visual tool use is learned from causally grounded supervision rather than tool-calling patterns.