发表机构
Shandong University; Nankai University; East China Normal University; National Technological University(山东大学; 南开大学; 华东师范大学; 国立科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对顺序诊断中平衡诊断准确性与资源成本的问题,提出GraphDx框架。该框架通过构建特殊知识图谱及引入协作智能体实现创新,实验表明其能显著提高诊断成功率并降低测试成本,为自动临床诊断提供有效方案。
AI 中文摘要
顺序诊断需要通过迭代信息收集来平衡诊断准确性和资源成本。现有的大语言模型(LLM)方法存在关键的知识推理差距:尽管编码了大量医学知识,但在成本约束下难以进行系统推理,常过度测试。我们提出GraphDx,一个有两个核心创新的知识增强框架。首先,设计自动化管道利用LLMs构建具有量化典型性、以行动为中心的拓扑结构以及诊断相关性和成本敏感性双目标属性的医学诊断知识图谱(MDKGs)。其次,引入三个协作智能体(感知、推理和决策),感知和决策智能体处理语言理解和生成,推理智能体在MDKG上进行确定性证据评分和成本感知规划。在MedQA和MIMIC-IV上针对三个LLM主干(DeepSeek-V3、Kimi-k2、Llama-3.3)的实验表明,GraphDx将诊断成功率从50%-68%提高到79%-93%,同时将测试成本降低20%-54%,为自动临床诊断提供了强大、经济且可解释的解决方案。
英文摘要
Sequential diagnosis requires balancing diagnostic accuracy against resource costs through iterative information gathering. Existing Large Language Model (LLM) approaches exhibit a critical knowledge-reasoning gap: despite encoding extensive medical knowledge, they struggle to reason systematically under cost constraints, often resorting to excessive testing. We propose GraphDx, a knowledge-enhanced framework with two core innovations. First, we design an automated pipeline that leverages LLMs to construct Medical Diagnosis Knowledge Graphs (MDKGs) with quantized typicality, action-centric topology, and dual-objective attributes for both diagnostic relevance and cost-sensitivity. Second, we introduce three collaborative agents (Perception, Reasoning, and Decision) where the Perception and Decision Agents handle language understanding and generation, while the Reasoning Agent performs deterministic evidence scoring and cost-aware planning on the MDKG. Experiments on MedQA and MIMIC-IV across three LLM backbones (DeepSeek-V3, Kimi-k2, Llama-3.3) show that GraphDx improves diagnostic success rates from 50--68% to 79--93% while reducing test costs by 20--54%, providing a robust, economical, and interpretable solution for automated clinical diagnosis.
Journal refFindings of the Association for Computational Linguistics: ACL 2026, pages 21721-21736, 2026
DOI:10.18653/v1/2026.findings-acl.1092