将医疗大语言模型(LLM)基于因果知识图谱:框架、指标与心血管试点研究
Framework for Grounding Healthcare LLMs in a Causal Knowledge Graph: A Cardiovascular Example Pilot
浏览论文内容
中文总结 AI 辅助
本研究提出以因果知识图谱为核心的医疗LLM评估框架,在心血管试点中验证其有效性,发现集成条件C4在因果推理相关指标上表现最优,未基于图的C1原始干预准确性最高但缺乏因果与证据基础。
中文摘要 AI 辅助
大语言模型(LLM)正越来越多地被提出用于医疗决策支持,但对它们的评估仍侧重于单一答案的准确性,而非对干预措施、机制、危害、证据及不确定性的推理。我们提出了一种可复现的、以图为中心的评估框架,用于医疗领域中面向干预措施的LLM行为,并在心血管试点中对其进行压力测试。该框架包含四个部分:(i)领域因果知识图谱,其中断言是具有稳定标识符的一等、保留来源的节点;(ii)场景条件子图提取步骤,给定任意临床场景,检索相关的具体化断言子图;(iii)四种受控的基于图的条件,用于改变检索到的子图如何组合到模型的上下文(未基于图的C1、知识图谱C2、因果图谱C3、集成C4);(iv)自动评分流程,以断言标识符为基准,计算干预准确性及其他单次评估指标。为测试该框架,我们构建了涵盖八种推理失败模式的类别平衡场景生成器,并在心血管图谱上实例化。指标面板沿可解释、非冗余轴区分各条件:C4获得最强的因果边F1(0.838)、不良反应F1(0.833)、证据准确性(0.738)及无依据主张率(0.114),而C1获得最高的原始干预准确性(0.948),但无可测量的因果或证据基础。
英文摘要
Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-centered evaluation framework for intervention-oriented LLM behavior in healthcare and stress-test it in a cardiovascular pilot. The framework has four components: (i) a domain causal knowledge graph in which assertions are first-class, provenance-preserving nodes with stable identifiers; (ii) a scenario-conditioned subgraph extraction step that, given any clinical scenario, retrieves the relevant reified-assertion subgraph; (iii) four controlled grounding conditions that vary how the retrieved subgraph is composed into the model's context (ungrounded C1, knowledge-graph C2, causal-graph C3, integrated C4); and (iv) an automated scoring pipeline, anchored on assertion identifiers, that computes intervention accuracy, and other evaluation measures on a single pass. To test the framework, we built a category-balanced scenario generator across eight reasoning failure modes and instantiated it on a cardiovascular graph. The metric panel discriminates conditions along interpretable, non-redundant axes: C4 obtains the strongest causal edge F1 (0.838), adverse-effect F1 (0.833), evidence accuracy (0.738), and unsupported claim rate (0.114), while C1 obtains the highest raw intervention accuracy (0.948) with no measurable causal or evidential grounding.