arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16131cs.CLcs.AI

ToolSciVer:基于视觉工具增强强化学习的多模态科学声明验证

ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning

Binglin Zhou, Peng Shi, Ryo Kamoi, Nan Zhang, Rui Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究多模态科学声明验证问题,提出ToolSciVer框架,为视觉语言模型配备三种类型感知视觉工具,用组相对策略优化训练策略,实验证明该方法在多模型上性能优于竞争基线。

中文摘要 AI 辅助

多模态科学声明验证(MSCV)要求模型利用来自论文的视觉证据(如图表、文本上下文)验证科学声明。现有方法常因难以定位关键视觉证据、准确读取结构化科学视觉内容及整合多模态观察进行可靠推理而失败。我们引入ToolSciVer,首个用于MSCV的工具增强框架。它为视觉语言模型配备三种类型感知视觉工具,通过组相对策略优化(GRPO)在复合奖励下训练策略。在三个模型家族的五个视觉语言模型上的实验表明,该方法性能优于四个竞争基线。

英文摘要

Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tables, charts, and textual context. However, existing methods often fail because they struggle to locate decisive visual evidence, accurately read structured scientific visuals, and integrate multimodal observations into reliable reasoning. We introduce ToolSciVer, the first tool-augmented framework for MSCV to our knowledge. ToolSciVer equips a VLM with three type-aware visual tools, table row/column focus, chart-to-structure parsing, and high-resolution region zoom, which convert dense scientific visuals into explicit, claim-facing evidence, and trains the policy with Group Relative Policy Optimization (GRPO) under a composite reward of answer correctness, format validity, length control, tool-use efficiency, and tool-validity penalties. Experiments on SciVer and MuSciClaims datasets on five VLMs from three model families (Qwen, InternVL, Gemma) demonstrate that our method achieves superior performance compared to four competitive baselines including prompting-based and RL-based tool-use methods, highlighting the effectiveness of learned, type-aware tool use for scientific claim verification.

发表机构

  • The Pennsylvania State University(宾夕法尼亚州立大学)
  • University of Waterloo(滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

↑