arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CDEG:为长程诊断智能体学习决策关键证据

CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents

Xiwei Dai, Zijie Meng, Zhiting Fan, Yixuan Tang, Guanyu Jiang, Ziru Niu, Zuozhu Liu

arXiv 2608.22899首次发表:更新:

AI 中文总结

本研究提出基于图的框架CDEG,通过对比同一病例的成功与失败轨迹、受控反事实干预等学习决策关键证据,在多基准测试中使长程诊断智能体准确率最高提升11.5%。

AI 中文摘要

与静态医学问答不同,长程诊断捕捉了临床实践的序列性:在达成最终诊断前,需经过多轮交互逐步获取、整合并评估证据。然而,现有医生智能体常因未获取关键证据或未将其充分纳入诊断推理而失效。近期的智能体方法试图通过复用历史轨迹或蒸馏记忆解决这些失效问题,但它们的诊断收益仍受限制,因为此类经验可能包含噪声或偶然信息,且通常在未验证哪些证据实际驱动诊断决策的情况下被复用。为解决这一局限,我们提出CDEG,一种基于图的框架,可从历史诊断轨迹中学习可复用的决策关键证据。CDEG对比同一病例的成功与失败轨迹以识别候选证据,通过受控反事实干预验证其诊断影响,并将所得的诊断-证据-动作关系组织为结构化图。推理阶段,CDEG跟踪不断演变的患者证据状态,以检索相关诊断关系,并选择性指导缺失证据的获取或被忽视证据的重新评估。在包含多种医生智能体主干的域内和域基准测试中,CDEG始终提升诊断性能,较普通智能体实现了高达11.5%的准确率提升。这些结果表明,可靠的长程诊断需超越轨迹级经验复用,转向对真正影响临床决策的因素进行证据级学习。

英文摘要

Unlike static medical question answering, long-horizon diagnosis captures the sequential nature of clinical practice: evidence is progressively acquired, integrated, and evaluated over multiple rounds of interaction before reaching a final diagnosis. However, existing doctor agents often fail when critical evidence is either not acquired or not adequately incorporated into diagnostic reasoning. Recent agentic approaches attempt to address these failures by reusing historical trajectories or distilled memories. But their diagnostic gains remain constrained because such experience may contain noisy or incidental information and is typically reused without validating which evidence actually drives diagnostic decisions. To address this limitation, we introduce CDEG, a graph-based framework that learns reusable decision-critical evidence from historical diagnostic trajectories. CDEG contrasts successful and failed trajectories from the same case to identify candidate evidence, validates their diagnostic impact through controlled counterfactual interventions, and organizes the resulting diagnosis--evidence--action relations into a structured graph. During inference, CDEG tracks the evolving patient evidence state to retrieve relevant diagnostic relations and selectively guide missing evidence acquisition or overlooked evidence reappraisal. Across in-domain and out-of-distribution benchmarks with multiple doctor agent backbones, CDEG consistently improves diagnostic performance, achieving up to an 11.5% accuracy gain over vanilla agents. These results demonstrate that reliable long-horizon diagnosis requires moving beyond trajectory-level experience reuse toward evidence-level learning of the factors that truly shape clinical decisions.

Comments9 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑