arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28225cs.CV

FaithEyes:通过多智能体过程图像验证实现可信的工具使用

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Self-Verification

Haoqing Wang, Xingrun Xing, Ziheng Li, Jianyuan Guo, Yehui Tang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出多智能体框架FaithEyes,通过子智能体判断主智能体的工具调用以提升工具使用可信性,在视觉感知与推理基准上实现了更优准确率。

中文摘要 AI 辅助

智能体视觉语言模型(VLMs)将文本推理与裁剪、基于代码的图像处理等显式工具调用相结合,已成为可靠且可解释的多模态推理的极具吸引力的范式。然而,近期研究表明这类模型常不可信地使用工具:许多过程图像与问题无关(例如工具裁剪了错误区域或遗漏了查询目标),但调用仍获得全额奖励,模型仍能正确回答。这种装饰性或不对齐的工具调用浪费了计算资源,表明模型依赖先验知识或原始图像而非其检索到的证据。这可能源于现有方法的两个局限:工具奖励无法区分有用与无用的调用,工具反馈不包含有用性信号。为此,我们引入FaithEyes,一个多智能体自判断框架。具体而言,我们使用一个VLM来判断每个过程图像是否有助于回答问题,该判断作为工具观察的一部分注入推理上下文以辅助后续推理,同时用于按有用工具比例缩放工具奖励,从而抑制奖励黑客行为。为使判断在评估阶段也可用,确保训练-测试一致性,我们进一步设计了多智能体框架,其中模型本身作为子智能体判断主智能体的工具调用,推理时无需依赖外部模型。通过在适配的开源数据上采用两阶段SFT+RL流程训练,FaithEyes在视觉感知和推理基准上达到了有竞争力或更优的准确率,同时显著提升了工具可信性。项目主页为this https URL。

英文摘要

Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, have emerged as a compelling paradigm for reliable and interpretable multi-modal reasoning. However, recent studies have revealed that such models often use tools unfaithfully. Many process images are irrelevant to the question (e.g., the crops miss the queried target), yet the tool call still receives full credit and the model still answers correctly. Such decorative or misaligned tool calls waste computation and reveal that the model does not faithfully use the evidence it retrieves. This may stem from two limitations of prevailing methods: the tool reward fails to distinguish useful from useless calls, and tool feedback carries no signal of usefulness. To this end, we introduce FaithEyes, a multi-agent self-judging framework. Concretely, we use a VLM to judge whether each process image helps answer the question. The judgement is injected into the reasoning context as part of the tool observation to help subsequent reasoning, and meanwhile is used to scale the tool reward by the helpful-tool ratio to suppress reward hacking. To keep judgement available at evaluation, we further design a multi-agent framework where the model itself serves as a subagent to judge the tool calls from the main agent, eliminating any dependence on external models at inference. Training via a two-stage SFT + RL pipeline on adapted open-source data, FaithEyes attains competitive or superior accuracy across visual perception and reasoning benchmarks, while substantially improving tool faithfulness and reducing inference cost. The homepage is at https://github.com/Mosi-AI/FaithEyes.

发表机构

  • Samsung Research(三星研究院)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

↑