引用你所探索的:基于可验证证据的预算感知LLM医学知识图谱推理
Cite What You Explore: Budget-Aware LLM Reasoning over Medical KGs with Verifiable Evidence
浏览论文内容
中文总结 AI 辅助
针对EHR出院后风险预测缺失依赖关系的问题,提出预算感知的LLM推理框架BAR,通过细化医学知识图谱为证据图并采用计划-导航-验证循环,在MIMIC数据集上显著提升AUPRC和引用精确率。
中文摘要 AI 辅助
从电子健康记录(EHRs)中进行出院后风险预测是困难的,因为许多将出院时观察结果与下游并发症(如合并症级联和药物-疾病相互作用)联系起来的依赖关系并未记录在病历中。外部医学知识图谱(KGs)可以提供这些缺失的依赖关系,但追溯这些关系需要满足三个特性:KG探索必须保持成本受限,检索到的证据必须按来源质量进行区分,并且最终得到的推理依据必须可被引用以供回顾性审查。大型语言模型(LLMs)能够对结构化证据进行规划和验证,使其成为KG推理的自然候选者,但现有的基于LLM的方法无法同时满足这三个特性。在本文中,我们提出了BAR,一个预算感知的LLM推理框架,用于医学知识图谱,包含三个贡献。首先,BAR将原始KG细化为疾病特定的证据图,其边携带支持分数和来源记录,将KG转变为质量注释的推理空间,而非静态特征源。其次,LLM通过一个计划-导航-验证循环在此图上进行推理,该循环将问题分解为步骤,在患者特定预算下检索证据,并在验证失败时进行修订。第三,训练一个推理策略,其奖励将有无获取证据时的预测进行比较,并结合获取成本和引用完整性项。在MIMIC-III和MIMIC-IV上,针对8种疾病和3个预测时间范围,BAR相比最强基线将AUPRC提高了3.4个百分点,将引用精确率从59.8%提升至77.9%,并且仅消耗预算上限的62-65%。
英文摘要
Post-discharge risk prediction from electronic health records (EHRs) is difficult because many dependencies that link discharge-time observations to downstream complications, such as comorbidity cascades and drug-disease interactions, are absent from the record. External medical knowledge graphs (KGs) can supply these missing dependencies, but tracing them demands three properties: KG exploration must remain cost-bounded, retrieved evidence must be differentiated by source quality, and the resulting rationale must be citable for retrospective review. Large language models (LLMs) can plan and verify over structured evidence, making them natural candidates for KG reasoning, but existing LLM-based methods do not satisfy these three properties jointly. In this paper, we propose BAR, a Budget-Aware LLM Reasoning framework over medical KGs with three contributions. First, BAR refines the raw KG into disease-specific evidence graphs whose edges carry support scores and provenance records, turning the KG into a quality-annotated reasoning space rather than a static feature source. Second, an LLM then reasons over this graph through a plan-navigate-verify loop that decomposes the question into steps, retrieves evidence under a patient-specific budget, and revises when verification fails. Third, a reasoning policy is trained with a reward that compares predictions with and without acquired evidence, combined with acquisition cost and citation-integrity terms. Across 8 diseases and 3 prediction horizons on MIMIC-III and MIMIC-IV, BAR improves AUPRC by 3.4 points over the strongest baseline, raises citation precision from 59.8% to 77.9%, and consumes only 62-65% of the budget cap.
发表机构
- University of Kansas(堪萨斯大学)
- University of Florida(佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。