arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30698cs.CV

MM-VeriAgent:通过强化学习学会使用多种工具验证多模态虚假信息

MM-VeriAgent: Learning to Use Extensive Tools to Verify Multimodal Misinformation with Reinforcement Learning

Peipei Li, Shuhan Xia, Shengyang Liu, Zekun Li, Ran He

首次发表
浏览论文内容

中文总结 AI 辅助

提出MM-VeriAgent,通过强化学习训练LVLM智能体使用工具包MM-VeriTools验证混合来源多模态虚假信息,并引入工具执行缓存提升训练效率,在MMFakeBench上取得显著准确率提升。

中文摘要 AI 辅助

现实世界的多模态虚假信息通常涉及混合伪造来源,需要针对样本特定的检测策略。现有的工具增强方法依赖于预定义的工作流程或推理时规划,这限制了适应性或增加了推理成本。为解决这一问题,我们提出了MM-VeriAgent,它学习使用工具来验证混合来源的多模态虚假信息。我们首先构建了MM-VeriTools,一个专门用于虚假信息检测智能体的工具包。通过在混合来源检测所需的子任务上对各种候选模型和方法进行基准测试,我们为文本、视觉和跨模态伪造分析选择了最强的模型,并将它们封装为具有统一接口的可调用工具。在此工具包的基础上,我们使用强化学习训练LVLM智能体,教会它如何使用这些工具来更好地解决混合来源检测问题。由于许多工具是专门模型,其在每次展开中的在线执行严重限制了强化学习效率,我们进一步引入了工具执行缓存,该缓存预执行候选工具调用并在训练期间重用其缓存输出。这保留了多步展开,同时减少了在线工具执行,大大提高了训练效率。在MMFakeBench上的实验结果表明,与基础模型相比,在推理时无需显式工具搜索即可获得显著的准确率提升。消融和效率分析进一步验证了所学的工具使用策略,并表明工具执行缓存减少了训练期间的在线工具执行次数。

英文摘要

Real-world multimodal misinformation often involves mixed forgery sources, requiring sample-specific detection strategies. Existing tool-augmented methods rely on predefined workflows or inference-time planning, limiting adaptability or increasing inference cost. To address this issue, we introduce \textbf{MM-VeriAgent}, which learns to verify mixed-source multimodal misinformation with tools. We first build \textbf{MM-VeriTools}, a specialized toolkit for misinformation detection agents. By benchmarking various candidate models and methods on the sub-tasks required by mixed-source detection, we select the strongest for textual, visual, and cross-modal forgery analysis and encapsulate them as callable tools with a unified interface. On top of this toolkit, we train the LVLM agent with reinforcement learning to teach it how to use these tools to better solve mixed-source detection. Since many of the tools are specialized models whose online execution at every rollout severely limits RL efficiency, we further introduce \textbf{Tool-Execution Cache}, which pre-executes candidate tool calls and reuses their cached outputs during training. This preserves multi-step rollouts while reducing online tool execution, largely improving the training efficiency.Experiments on MMFakeBench demonstrate substantial accuracy gains over the base model without explicit tool search at inference time. Ablation and efficiency analyses further validate the learned tool-use policy and show that Tool-Execution Cache reduces online tool executions during training.

发表机构

  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • Minzu University of China(中央民族大学)
  • NLPR, Institute of automation, Chinese academy of science, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑