发表机构
Institute of Information Engineering, Chinese Academy of Sciences; Department of Machine Learning, Carnegie Mellon University; MiLM Plus, Xiaomi Inc.(中国科学院信息工程研究所; 卡内基梅隆大学机器学习系; 小米公司MiLM Plus)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体视觉语言模型工具调用监督不足的问题,提出NTEP标注方案与NTEP-R监督机制,在7个基准上验证了NTEP-8B可提升搜索准确率与工具使用效率。
AI 中文摘要
现代视觉语言模型(VLMs)可直接回答许多基于图像的问题,但在需要细粒度视觉细节或外部知识的复杂查询上常表现不佳。为获取缺失的证据,智能体视觉语言模型会调用图像裁剪、图像搜索、文本搜索等工具。然而,现有训练范式主要基于最终答案正确性评估工具使用,对证据获取和利用的监督不足,导致两个关键缺陷:(i)模型常发出冗余或偏离目标的工具调用,无法收集必要证据;(ii)即使调用了合适的工具,模型也常无法从结果观测中提取必要信息。为解决这些局限,我们提出NTEP(必要工具-证据路径),一种为每个查询明确指定必要外部证据及对应工具调用的新型标注方案。在此基础上,我们提出NTEP-R(NTEP奖励),一种确保每次工具调用严格推进推理过程至最终解决方案的监督机制。具体而言,该方法奖励智能体在调用前将意图与必要证据寻求目标对齐,以及确保从调用后观测中总结的信息与必要证据对齐。此外,我们引入非重复目标正则化项,惩罚重复访问已满足的NTEP目标的冗余调用。在7个基于图像的基准上的广泛评估表明,我们的8B参数实例NTEP-8B在统一的三工具框架内,显著提升了面向搜索的准确率和工具使用效率。这些结果凸显了细粒度工具-证据路径监督对训练鲁棒智能体视觉语言模型的关键价值。
英文摘要
Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or external knowledge. To acquire this missing evidence, agentic VLMs invoke tools such as image cropping, image search, and text search. However, existing training paradigms primarily evaluate tool-use based on final answer correctness, leaving evidence acquisition and utilization insufficiently supervised. This leads to two critical shortcomings: (i) models frequently issue redundant or off-target tool calls that fail to gather necessary evidence, and (ii) even when appropriate tools are called, models often fail to extract the necessary information from the resulting observations. To address these limitations, we introduce the NTEP (Necessary Tool-Evidence Path), a novel annotation scheme that explicitly specifies the essential external evidence and corresponding tool calls for each query. Building upon this, we propose NTEP-R (NTEP Reward), a supervision mechanism ensuring that each tool invocation strictly advances the reasoning process toward the final solution. Specifically, our approach rewards the agent for aligning its pre-call intent with a necessary evidence-seeking goal, and for ensuring the information summarized from the post-call observation aligns with the necessary evidence. Furthermore, we introduce a non-repeated-goal regularizer to penalize redundant calls that revisit satisfied NTEP goals. Extensive evaluations on seven image-grounded benchmarks demonstrate that our 8B-parameter instantiation, NTEP-8B, significantly improves both search-oriented accuracy and tool-use efficiency within a unified three-tool framework. These results highlight the critical value of fine-grained tool-evidence path supervision for training robust agentic VLMs.