arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CRAFT:基于大语言模型的迭代优化用于临床叙事的时间推理

CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives

Chengyang He, Tahreem Arif, Marko Zivkovic, Lijing Wang, Yue Ning, Ping Wang

arXiv 2608.12779首次发表:更新:

发表机构

Stevens Institute of Technology; Genesis Research Group; New Jersey Institute of Technology(史蒂文斯理工学院; 创世纪研究集团; 新泽西理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对临床叙事中时间锚点稀疏导致的症状轨迹重建难题,提出基于大语言模型的CRAFT框架,通过生成器与约束验证器迭代优化,在MedTempo基准上提升了时间排序准确率。

AI 中文摘要

理解临床叙事中症状的时间进展对疾病监测、安全监测和因果关系评估至关重要。然而,临床叙事很少提供明确的时间锚点。当前的时间信息推理方法主要关注多就诊、时间戳丰富的记录中的成对关系分类,而从单个锚点稀疏的报告中重建结构化症状轨迹的问题在很大程度上未得到解决。我们提出CRAFT,这是一个大语言模型框架,它将生成器与基于约束的验证器配对,通过针对性反馈迭代生成和优化分阶段症状时间线。我们在MedTempo上进行评估,MedTempo是一个新的基准,包含5347篇疫苗不良事件叙事,涵盖三种新冠疫苗类型,其中3166篇报告具有专家验证的时间阶段注释。对四种大语言模型主干的实验表明,CRAFT始终提高了时间排序准确率,而 ablation分析则在不同模型能力水平上分离了生成器和验证器组件的贡献。

英文摘要

Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment. Clinical narratives, however, rarely provide explicit temporal anchors. Current approaches to temporal information reasoning focus predominantly on pairwise relation classification across multi-visit and timestamp-rich records, leaving the reconstruction of structured symptom trajectories from individual anchor-sparse reports largely unaddressed. We propose CRAFT, an LLM framework that pairs a generator with a constraint-based verifier to iteratively produce and refine stage-wise symptom timelines through targeted feedback. We conduct evaluation on MedTempo, a new benchmark of 5,347 vaccine adverse-event narratives spanning three COVID-19 vaccine types, with expert-validated temporal stage annotations for 3,166 reports. Experiments across four LLM backbones demonstrate that CRAFT consistently improves temporal ordering accuracy, with ablation analysis isolating the contribution of generator and verifier components across model capability levels.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑