arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.27860cs.AI

C-MIG:基于多视角信息增益的检索增强生成用于临床诊断推理

C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning

  • Baidu Inc(百度公司)

机构由 AI 辅助整理,请以论文原文为准。

Yuwei Miao, Gen Li, Yunsheng Zeng, Xiandong Li, Yujin Wang, Siyu Chen, Luning Wang, Yunhao Qiao, Junfeng Wang, Jianwei Lv, Bo Yuan

AI总结:

提出C-MIG框架,通过多视角信息增益和多重子查询检索增强策略,解决检索增强生成中奖励信号丢失和异构推理监督问题,在临床诊断任务上取得最优性能。

AI中文摘要:

检索增强生成结合强化学习在将大型语言模型锚定于可信医学证据方面显示出前景。然而,现有方法依赖精确匹配的二元奖励,在临床诊断中导致两个问题:(i) 语义相关但非逐字匹配的步骤获得零信号,丢弃了有价值的学习信号;(ii) 单一维度的奖励无法有效监督异构推理能力。为解决这些问题,我们提出C-MIG,一种基于多视角信息增益的临床诊断检索增强生成框架。C-MIG在冻结参考模型下从两个互补视角——检索文档和文档精炼——估计信息增益,以联合指导检索什么以及如何精炼,缓解了有价值奖励信号丢失和信用分配问题。我们进一步设计了一种多重子查询检索增强策略,提高了临床诊断场景中的知识召回覆盖率。在四个医学基准上的综合实验表明,C-MIG在领域内和领域外数据集上均达到所有RAG-RL方法中的最佳性能,并在临床诊断上超越了最先进的通用大型语言模型。

英文摘要:

Retrieval-augmented generation combined with reinforcement learning has shown promise for grounding large language models in trustworthy medical evidence. However, existing methods rely on exact-match binary rewards, which in clinical diagnosis cause two issues: (i) semantically relevant but non-verbatim steps receive zero signal, discarding valuable learning signals; and (ii) uni-dimensional rewards cannot effectively supervise heterogeneous reasoning capabilities. To address these issues, we propose C-MIG, a Multi-view Information Gain-based retrieval-augmented generation framework for Clinical diagnosis. C-MIG estimates information gain under a frozen reference model from two complementary views, retrieved-document and document-refinement, to jointly guide what to retrieve and how to refine, alleviating the issues of valuable reward signal loss and credit assignment. We further design a multi-subquery retrieval augmentation strategy that improves knowledge recall coverage in clinical diagnostic scenarios. Comprehensive experiments on four medical benchmarks demonstrate that C-MIG achieves the best performance among all RAG-RL methods on both in-domain and out-of-domain sets, and outperforms state-of-the-art general-purpose LLMs for clinical diagnosis.

↑