发表机构
Bordeaux University Hospital (CHU Bordeaux)(波尔多大学医院(CHU Bordeaux))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出基于溯源感知证据图的可审计临床时间线重建方法,通过保留提及、记录修订及弃权(不执行)机制,在合成语料上验证了可审计性,并评估了BERT与LLM的证据定位能力。
AI 中文摘要
患者时间线重建系统只有在保留每个答案背后的提及、记录事实如何被修订、并在证据不在文本中时拒绝回答的情况下,才是可审计的。本研究在一个完全合成的语料库(1000名患者,3353条笔记,220条修订边)上测试了这三个属性。两种溯源感知证据图算子将节点加边的数量减少到67%和63%(序列化大小的77-78%),同时保留了6813个可由最近性回答的查询点上的每个答案和提及链接;一个固定窗口基线在53.4%的点上未返回任何值,且未标记。在宣布遗漏的证据不可用对照中,BioClinicalBERT门和零样本LLM门主要对公告作出响应。在无标记对照中,BERT在81条笔记中的0条上弃权(不执行),而其准确性在所有三个关系类别中从93.8%下降到59.3%;LLM的覆盖率在其自身模型家族判定为无法确定的笔记上从75.6%下降到27.7%。针对482个再生成的金标准片段,LLM引用的证据达到了0.850的召回率和0.864的精确率;BERT的跨度头在没有跨度标签训练的情况下,未能定位证据。一个时间版本化的溯源图将弃权(不执行)存储为类型化、可查询的边。该干净任务允许0.651准确率的捷径,结果描述了在合成数据上的实现行为,而非临床性能。
英文摘要
A patient-timeline reconstruction system is auditable only if it keeps the mentions behind each answer, records how facts were revised, and declines to answer when the evidence is not in the text. This study tests these three properties on a fully synthetic corpus (1,000 patients, 3,353 notes, 220 revision edges). Two provenance-aware Evidence Graph operators reduced the node-plus-edge count to 67% and 63% (77-78% of serialized size) while preserving every answer and mention link across 6,813 query points answerable by recency; a fixed-window baseline returned no value for 53.4% of points, unflagged. On evidence-unavailable controls that announce the omission, a BioClinicalBERT gate and a zero-shot LLM gate responded mainly to the announcement. On marker-free controls, BERT abstained on 0 of 81 notes while its accuracy fell from 93.8% to 59.3% across all three relation classes; the LLM's coverage fell from 75.6% to 27.7% on notes its own model family judged undeterminable. Against 482 regenerated gold spans, the LLM's cited evidence reached recall 0.850 and precision 0.864; BERT's span head, trained without span labels, did not localize evidence. A temporally versioned provenance graph stored abstentions as typed, queryable edges. The clean task admits a 0.651-accuracy shortcut, and results describe implementation behaviour on synthetic data, not clinical performance.
Comments30 pages, 11 Tables, 6 figures