发表机构
Beijing University of Posts and Telecommunications; Peking University; WeChat AI, Tencent Inc.(北京邮电大学; 北京大学; 腾讯微信人工智能部门)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有无预设问题/证据的学术论文科学错误验证研究不足的问题,提出VERA-RL强化学习框架,构建VERA-13K数据集,训练Qwen3-VL-8B后其可验证推理能力接近旗舰级多模态大语言模型。
AI 中文摘要
多模态大语言模型(MLLMs)正日益成为强大的科学助手,但远未实现完全自主的研究。这一转变要求模型主动检查学术论文、构建全局证据视图,并在无预设问题或证据的情况下做出可追溯的判断。然而,现有工作针对这种无问题、无证据的验证任务范式或训练研究有限。我们以科学错误检测为研究对象,要求模型判断是否存在错误并通过基于证据的推理进行论证。为填补这一空白,我们提出VERA-RL,一种面向学术论文科学错误检测的强化学习框架。遵循“推理-验证-扫描”流程,我们构建了VERA-13K,包含12900个样本,分为4300个匹配链,涵盖研究流程中的6类科学错误及广泛的自然科学领域。我们进一步引入细粒度奖励,用于推理完整性、证据对齐及错误精度。使用VERA-RL训练Qwen3-VL-8B可显著提升可验证推理能力,在扫描任务上接近Gemini 3 Pro、Qwen3-VL-235B-A22B等旗舰级多模态大语言模型。
英文摘要
Multimodal large language models (MLLMs) are increasingly capable scientific assistants, yet they remain far from fully autonomous research. This transition requires models to actively inspect academic papers, build global evidence views, and make traceable judgments without prespecified issues or evidence. However, existing work provides limited task paradigms or training studies for such issue- and evidence-absent verification. We study this challenge through scientific error detection, where models must determine whether errors exist and justify them with evidence-based reasoning. To fill this gap, we present VERA-RL, a reinforcement-learning formulation for scientific error detection over academic papers. Following a Reason--Verify--Scan progression, we construct VERA-13K, a 12,900-sample dataset organized into 4,300 matched chains, covering 6 scientific-error categories across the research workflow and broad natural-science domains. We further introduce fine-grained rewards for reasoning completeness, evidence alignment, and error precision. Training Qwen3-VL-8B with VERA-RL substantially improves verifiable reasoning, approaching flagship MLLMs such as Gemini 3 Pro and Qwen3-VL-235B-A22B on Scan.
CommentsAccepted by EMNLP 2026 Findings