arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19585cs.CLcs.AIcs.IRcs.LG

CliniCIRCA:一种从原始EHR叙事构建纵向心理健康患者轨迹的模块化LLM框架

CliniCIRCA: A Modular LLM Framework for Constructing Longitudinal Mental Health Patient Journeys from Raw EHR Narratives

  • Georgia Institute of Technology(佐治亚理工学院)
  • University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • Northwell Health(诺斯韦尔健康)

机构由 AI 辅助整理,请以论文原文为准。

Aiwei Ivy Zhang, Nimra Ishfaq, Mohit Chandra, Santiago Alvarez Lesmes, Adam Coscia, Khatiya Chelidze Moon, Xiaohan Ding, Munmun De Choudhury

AI总结:

CliniCIRCA是一个多阶段LLM框架,用于从非结构化EHR叙事中重建纵向心理健康患者轨迹,通过时间分类和总结,生成时间线并提升事件提取与总结性能。

AI中文摘要:

在心理健康护理中,推理患者轨迹是临床医生的关键任务。然而,这些轨迹涵盖了生物、心理和社会事件的纵向进展,通常分散在不同的非结构化文本叙述中,使得时间恢复具有挑战性。我们提出了CliniCIRCA,一个用于日历锚定、不精确感知的临床年鉴重建的多阶段LLM框架。据我们所知,CliniCIRCA是首个在没有事件级时间戳的情况下对非结构化出院总结中的临床事件进行时间分类的框架。从14,882例MIMIC-III心理健康入院记录中,我们首先构建了一个包含52份出院总结的基准,CliniCIRCA在其上生成了15,891个带时间标记的事件。在基于临床医生参与评估纠正629个错误后,我们产生了经过验证的金标准标签。最后,纠正后的时间线驱动一个时间接地总结阶段,将每个源文档压缩1.52倍,形成按日期分组的时间顺序记录。然后,我们将框架扩展以生成1,000条银标准时间线,并评估它们作为训练数据。与零样本和少样本提示相比,指令微调通常在五个开放权重模型上改善了事件提取、时间标记和总结,在银标准和临床医生验证的评估中均表现更好。

英文摘要:

In mental health care, reasoning over patient journeys is a key task for clinicians. Yet these journeys, encompassing a longitudinal progression of biological, psychological, and social events, are often spread across disparate unstructured text narratives, making temporal recovery challenging. We present CliniCIRCA, a multi-stage LLM framework for Calendar-anchored, Imprecision-aware Reconstruction of Clinical Annals. To our knowledge, CliniCIRCA is the first to temporally classify clinical events across unstructured discharge summaries without event-level timestamps. From 14,882 MIMIC-III mental health admissions, we first construct a benchmark of 52 discharge summaries on which CliniCIRCA produces 15,891 temporally tagged events. After correcting 629 errors based on a clinician-in-the-loop evaluation, we produce verified gold-standard labels. Finally, the corrected timelines drive a temporally grounded summarization stage that compresses each source 1.52 times into a date-grouped chronological record. We then scale the framework to generate 1,000 silver-standard timelines and evaluate them as training data. Compared with zero- and few-shot prompting, instruction tuning generally improves five open-weight models on event extraction, temporal tagging, and summarization across silver and clinician-verified evaluations.

↑