arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.25566cs.AI

基于大语言模型的不确定性推理用于可解释疾病诊断

Uncertainty Reasoning with Large Language Models for Explainable Disease Diagnosis

  • National University of Singapore(新加坡国立大学)
  • Griffith University(格里菲斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiaoyang Fan, Yufan Cai, Zhe Hou, Jin Song Dong

更新

AI总结:

提出一种神经符号推理框架,将大语言模型与模糊逻辑和声明式规则结合,实现可解释且形式可验证的医学诊断。

AI中文摘要:

临床决策需要对不完整、不精确且以语言表达的患者叙述进行推理。虽然大语言模型(LLMs)擅长从自然语言中提取潜在信息,但它们缺乏可信赖医疗AI所必需的可验证性和可解释性。我们提出一种神经符号推理框架,将LLMs与形式逻辑对齐,以实现可解释且形式可验证的医学诊断。患者描述和临床指南被嵌入神经知识库,其中LLMs提取结构化医疗实体、时间关系和模糊症状模式,这些被解码为用模糊逻辑和声明式规则表达的符号知识库。我们执行两阶段推理:(1)归纳符号泛化,从编码叙述中捕获诊断模式;(2)通过逻辑编程引擎进行推理验证,推导并验证符合临床标准的诊断。每个症状被视为具有概率权重的模糊谓词,推理路径可审计、可调整,并与医生反馈兼容。与纯统计方法不同,我们的系统支持迭代优化:LLM生成的诊断与真实情况之间的偏差可以通过形式规则追踪、解释和纠正。通过结合基于逻辑的透明性、LLM的适应性和概率鲁棒性,该框架实现了与人类一致的医疗推理,具有强泛化能力和可验证的逐步推理链。我们在公开基准上验证了该框架,展示了符号推理与LLM在真实临床叙述中的有效协调。结果显示,性能与最先进的LLM相当,同时额外提供了可解释的推理路径和形式可验证的诊断结论。

英文摘要:

Clinical decision-making requires reasoning over incomplete, imprecise, and linguistically expressed patient narratives. While large language models (LLMs) excel at extracting latent information from natural language, they lack the verifiability and interpretability essential for trustworthy medical AI. We propose a neuro-symbolic reasoning framework that aligns LLMs with formal logic to enable explainable and formally verifiable medical diagnosis. Patient descriptions and clinical guidelines are embedded into a neural knowledge base, where LLMs extract structured medical entities, temporal relations, and fuzzy symptom patterns, which are decoded into a symbolic knowledge base expressed in fuzzy logic and declarative rules. We perform two-stage reasoning: (1) inductive symbolic generalization to capture diagnostic patterns from encoded narratives, and (2) inference verification via a logic programming engine to derive and validate diagnoses consistent with clinical standards. Each symptom is treated as a fuzzy predicate with probabilistic weights, and inference paths are auditable, adjustable, and compatible with physician feedback. Unlike purely statistical methods, our system supports iterative refinement: misalignment between LLM-generated diagnoses and ground truth can be traced, explained, and corrected through formal rules. By combining logic-based transparency, LLM adaptability, and probabilistic robustness, the framework enables human-aligned healthcare inference with strong generalization and verifiable, step-by-step reasoning chains. We validate our framework on public benchmarks, demonstrating effective reconciliation of symbolic reasoning and LLMs with real-world clinical narratives. Results show performance comparable to state-of-the-art LLMs, while additionally providing interpretable reasoning paths and formally verifiable diagnostic conclusions.

↑