arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23811cs.AI

用强化学习增强的智能体搜索生成生物医学事实核查报告

Generating Biomedical Fact-Checking Reports with RL-Enhanced Agentic Search

Jiongxiao Wang, Dingli Ma, Chaoqun Ni

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出BioCheck Agent智能体结合EG-GRPO强化学习方法,生成生物医学事实核查报告,在SciFact数据集上提升了标签预测准确率、证据质量并降低了幻觉率。

中文摘要 AI 辅助

自动化事实核查对于确保公共卫生信息的可靠性至关重要,但生物医学领域存在独特挑战。验证生物医学主张需要对科学文献进行严格解读、评估检索到的证据,并为结论提供全面依据。尽管检索增强生成(RAG)和智能体搜索增强的大语言模型(LLM)以“先检索后验证”的范式执行自动化事实核查,但现有方法仍输出孤立的预测标签,缺乏解释深度,对人类理解的效用有限。为填补这一空白,我们引入了名为BioCheck Agent的基于LLM的智能体,该智能体通过智能体搜索生成结构化生物医学事实核查报告。BioCheck Agent不仅输出支持或反驳的标签,还会结合检索到的证据与严格分析得出最终结论。为确保领域特定的准确性,BioCheck Agent仅搜索PubMed中的高质量科学文献,利用高级布尔搜索运算符。考虑到直接提示往往会导致幻觉和低质量报告,尤其是对于轻量级开源模型,我们进一步提出了证据基础组相对策略优化(EG-GRPO),以针对BioCheck Agent执行强化学习,采用特定任务的奖励机制,激励高级搜索行为和高质量证据检索,同时惩罚幻觉。实验结果表明,与基础模型Qwen3.5-4B相比,采用EG-GRPO的BioCheck Agent在SciFact上的标签预测准确率提高了9.95%;此外,其证据质量得分提高了3.7%,证据幻觉率降低了19.63%,证明了其生成准确性和质量均有所提升的生物医学事实核查报告的能力。

英文摘要

Automated fact-checking is essential for ensuring the reliability of public health information, yet the biomedical domain poses unique challenges. Validating biomedical claims requires rigorous interpretation of scientific literature, assessment of retrieved evidence, and comprehensive justification toward the conclusion. Although Large Language Models (LLMs) enhanced by Retrieval-Augmented Generation (RAG) and agentic search perform automated fact-checking in a retrieve-then-verify paradigm, current methods still output isolated prediction labels, lacking explanatory depth and offers limited utility for human understanding. To bridge this gap, we introduce an LLM-based agent named BioCheck Agent that generates structured biomedical fact-checking reports with agentic search. Rather than merely outputting supported or refuted labels, our agent synthesizes final conclusions with retrieved evidence and rigorous analysis. To ensure domain-specific accuracy, BioCheck Agent exclusively searches high-quality scientific literature in PubMed, utilizing advanced Boolean search operators. Recognizing that direct prompting often results in hallucinations and low-quality reports, especially for lightweight open-source models, we further propose the Evidence-Grounded Group Relative Policy Optimization (EG-GRPO) to perform reinforcement learning on BioCheck Agent with a task-specific reward that incentivizes advanced search behavior and high-quality evidence retrieval while penalizing hallucinations. Our experimental results show that compared to the base model Qwen3.5-4B, BioCheck Agent with EG-GRPO improves label prediction accuracy on SciFact by 9.95%. Furthermore, it achieves a 3.7% higher evidence quality score and a 19.63% lower evidence hallucination rate, demonstrating its ability to generate biomedical fact-checking reports with improved accuracy and quality.

发表机构

  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

↑