VC-Tooler:学习组合式与自适应视觉工具使用
VC-Tooler: Learning Compositional and Adaptive Visual Tool Use
AI总结:
VC-Tooler是一种学习组合式与自适应视觉工具使用的模型,通过分层合成轨迹库与两阶段训练,在通用和具身基准上达到开源模型最优性能,且具良好迁移能力。
AI中文摘要:
具身多模态推理通过让视觉语言模型(VLMs)借助视觉工具交互主动获取和细化视觉证据,拓展了被动图像理解。有效的视觉工具使用需要三项能力:将工具调用基于视觉语境、跨多步骤组合工具、根据工具返回的观测结果调整推理。然而现有方法大多聚焦于固定工具空间内的基础调用模式,对组合性与适应性的处理不足。本文提出VC-Tooler,将视觉工具使用作为组合式与自适应能力进行学习。为此,我们首先通过分层合成管线构建轨迹库,涵盖单工具基础、多工具组合、多样工具语境与接口三个能力层级;随后分两阶段训练模型:建立上述能力的监督冷启动阶段,以及鼓励准确、高效、语境感知的视觉工具使用的强化学习阶段。VC-Tooler在通用与具身基准的开源模型中达到最优性能,包括V*上的95.8%和VTC-Bench上的35.3%,且在推理时更丰富的工具设置下展现出良好的迁移能力。项目页面:this https URL
英文摘要:
Agentic multimodal reasoning extends passive image understanding by allowing VLMs to actively acquire and refine visual evidence through visual tool interactions. Effective visual tool use requires three capabilities: grounding tool calls in visual context, composing tools across multiple steps, and adapting reasoning to tool-returned observations. However, existing approaches largely focus on grounding within fixed tool spaces and rigid invocation patterns, leaving composition and adaptation insufficiently addressed. We present VC-Tooler, which learns visual tool use as a compositional and adaptive capability. To this end, we first build a trajectory bank through a hierarchical synthesis pipeline covering three capability levels: single-tool grounding, multi-tool composition, and diverse tool contexts and interfaces. We then train the model in two stages: a supervised cold start that establishes these capabilities, followed by reinforcement learning that encourages accurate, efficient, and context-aware visual tool use. VC-Tooler achieves state-of-the-art performance among open-source models on both general-purpose and agentic benchmarks, including $95.8\%$ on V* and $35.3\%$ on VTC-Bench, and shows promising transfer under richer tool settings at inference time. Project page: https://w1zheng.github.io/VC-Tooler