AI 中文总结
本研究提出MedEventGraph-RAG框架,通过图结构关联患者临床事件与源证据,兼顾正反证据检索与过滤,在四个数据集的纵向临床事件验证任务中显著提升性能,减少无依据结论。
AI 中文摘要
纵向临床事件关系验证旨在判断患者记录是否支持两个或多个临床事件间的指定关系。该任务极具挑战性,因为证据分散在结构化记录、病历文本、实验室轨迹、就诊记录及时间维度中,而否定表述、时间不匹配、重复记录及相互矛盾的结果可能导致检索信息看似相关却无法确立目标关系。本文提出MedEventGraph-RAG,这是一种证据可接纳框架,它将患者特有的临床事件发生情况表示为图结构,并将每个事件发生情况与源证据关联,源证据包括结构化行、病历文本片段、时间戳及数值轨迹。给定指定事件、关系及临床范围的验证查询后,该图会引导发现候选事件链,并从支持与矛盾两方检索证据。在单独评估者判断支持、矛盾、被反驳或证据不足的结果前,会先通过查询特定的证据契约,按患者身份、范围、事件绑定及源可追溯性对信息进行过滤。在i2b2、n2c2、MIMIC-IV及LUNGUAGE四个数据集的十个协议实验中,MedEventGraph-RAG在时间验证、药物不良事件验证及医嘱记录验证任务上的平衡准确率分别达到78.6、67.3和96.8,较最强匹配基线分别提升26.9、4.9和30.4个百分点;在证据遮蔽场景下,其平衡准确率达92.2且无错误支持预测;当中间事件被隐藏时,它在i2b2数据集中57.9%的案例、LUNGUAGE数据集中70.0%的案例中可恢复完整的源可追溯事件链。这些结果表明,将广泛的证据发现与精准的证据可接纳评估分离,可提升纵向临床验证性能并减少无依据结论。
英文摘要
Longitudinal clinical event-relation verification determines whether a patient record supports a specified relation among two or more clinical events. This task is challenging because evidence is distributed across structured records, notes, laboratory trajectories, encounters, and time, while negation, temporal mismatch, repeated documentation, and conflicting findings can make retrieved information appear relevant without establishing the relation. We present MedEventGraph-RAG, an evidence-admissible framework that represents event occurrences in a patient-specific graph and links each occurrence to source evidence, including structured rows, note spans, timestamps, and numerical trajectories. Given a verification query specifying events, relation, and clinical scope, the graph guides discovery of candidate event chains and retrieves evidence from both supporting and contradicting sides. A query-specific evidence contract filters information by patient identity, scope, occurrence binding, and source traceability before a separate assessor determines supported, conflicting, refuted, or insufficient outcomes. Across ten protocols on i2b2, n2c2, MIMIC-IV, and LUNGUAGE, MedEventGraph-RAG achieves balanced accuracies of 78.6, 67.3, and 96.8 on temporal, medication-adverse-event, and recorded-order verification, improving over the strongest matched baselines by 26.9, 4.9, and 30.4 points. Under evidence masking, it reaches 92.2 balanced accuracy with no false-support predictions. When intermediate events are hidden, it recovers complete source-traceable event chains in 57.9% of i2b2 and 70.0% of LUNGUAGE cases. These results show that separating broad evidence discovery from narrow evidence-admissible assessment improves longitudinal clinical verification and reduces unsupported conclusions.
CommentsSubmitted to AAAI 2027