MedAgent-R1:面向证据 grounded 医疗推理的忠实性感知强化学习
MedAgent-R1: Faithfulness-Aware Reinforcement Learning for Evidence-Grounded Medical Reasoning
浏览论文内容
中文总结 AI 辅助
MedAgent-R1 针对医疗推理智能体的自信幻觉问题,采用忠实性门控奖励设计,大幅降低引用编造率、提升证据完整性,在 HealthBench Safety 上表现优于 GPT-4o,实现了证据 grounded 的医疗推理优化。
中文摘要 AI 辅助
当医疗 AI 系统在临床推理中产生幻觉时,其后果不仅限于答案错误:编造的理由虽表面上引用了检索到的证据,却可能误导临床医生做出不安全的治疗决策。因此,医疗推理智能体不仅必须给出正确答案,还必须提供临床医生可对照引用证据进行验证的忠实理由。我们发现了强化学习(RL)训练的检索智能体中的一种系统性失效模式:仅基于结果的奖励会提升准确性,却会降低忠实性,我们将这一现象称为“自信幻觉”。该智能体学会从参数化记忆中获取答案,并补充看似合理但无依据的理由;尽管准确率较监督基线提升了 5 个百分点,引用编造率却从 16.5% 升至 31.8%。我们通过一种忠实性门控奖励设计解决了这一问题:准确性奖励需通过硬门控(hard gate)以证据依据为条件,同时辅以检索有效性和简洁性信号,以关闭智能体检索特有的利用路径。由此得到的系统 MedAgent-R1 将引用编造率从 31.8% 降至 4.7%,证据完整性从 58.7 提升至 82.6,同时保持 75.1% 的准确率,在 HealthBench Safety 上的得分提升了 13.2 个百分点。在相同的智能体检索设置下,MedAgent-R1 在忠实性特定维度上的得分超过 GPT-4o(事实支持维度:4.55 对 4.25;过度主张维度:4.40 对 4.15),尽管整体准确率仍低于 GPT-4o,这表明显式忠实性训练可带来仅通过扩大模型规模无法实现的证据依据提升。
英文摘要
When medical AI systems hallucinate clinical reasoning, the consequences extend beyond incorrect answers: fabricated justifications that superficially reference retrieved evidence can mislead clinicians into unsafe treatment decisions. Medical reasoning agents must therefore produce not only correct answers but also faithful justifications that clinicians can verify against cited evidence. We identify a systematic failure mode in RL-trained retrieval agents: outcome-only rewards improve accuracy while degrading faithfulness, a phenomenon we term confident hallucination. The agent learns to answer from parametric memory and backfill plausible but unsupported justifications; citation fabrication rates rise from 16.5% to 31.8% even as accuracy improves by 5 points over the supervised baseline. We address this with a faithfulness-gated reward design: accuracy credit is conditioned on evidence grounding via a hard gate, complemented by retrieval validity and conciseness signals that close exploitation paths unique to agentic retrieval. The resulting system, MedAgent-R1, reduces citation fabrication from 31.8% to 4.7% and raises evidence completeness from 58.7 to 82.6 while maintaining 75.1% accuracy, with 13.2-point gains on HealthBench Safety. Under the same agentic retrieval setup, MedAgent-R1 outscores GPT-4o on faithfulness-specific dimensions (Factual Support 4.55 vs. 4.25; Overclaiming 4.40 vs. 4.15) while remaining below GPT-4o in overall accuracy, suggesting that explicit faithfulness training yields evidence-grounding gains not achieved by scaling alone.
发表机构
- Tsinghua University(清华大学)
- DP Technology(第四范式)
机构由 AI 辅助整理,请以论文原文为准。