发表机构
University of Georgia; Beijing Luhe Hospital(佐治亚大学; 北京潞河医院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出EarlyDx基准,针对现有诊断基准不适用入院诊断场景的问题,基于MIMIC-IV数据构建,发现各类LLM均无法可靠综合入院证据,仅能提取部分诊断,微调后仍有差距。
AI 中文摘要
医院入院时的临床诊断必须从有限、不完整的证据中快速做出。现有的诊断预测基准不适合这种场景:它们将预测限制在封闭代码集,排除自由文本记录,并使用包含整个住院过程的出院诊断作为监督。我们引入了EarlyDx,这是一个基于MIMIC-IV中154834次急诊科就诊构建的大规模开放式早期诊断基准。每次就诊仅限于入院时间t₀可用的记录,并以急诊科就诊期间记录的诊断而非出院诊断作为监督。一个大语言模型(LLM)审核员进一步验证每个自由文本标签是否得到该证据的支持、部分支持或不支持;主要评估仅对完全支持的标签打分。在语义LLM作为评判者的协议下,所有评估的系统——前沿通用模型、医学专业模型或领域内微调模型——都无法可靠地综合入院时间的证据。零样本模型主要通过提取得分,仅恢复3%-31%必须推断而非从记录中直接读取的诊断;微调将依赖推断的召回率提升至56%,但仍存在较大差距,且对于时间关键型病症,没有任何系统达到临床医生的敏感性与精确性平衡。我们在此发布完整的构建与评估流程。
英文摘要
Clinical diagnosis at hospital admission must be made rapidly from limited, incomplete evidence. Existing diagnosis-prediction benchmarks are poorly suited to this setting: they restrict prediction to closed code sets, exclude free-text notes, and supervise with discharge diagnoses that incorporate the full inpatient course. We introduce EarlyDx, a large-scale benchmark for open-ended early diagnosis, built from 154,834 emergency department encounters in MIMIC-IV. Each encounter is restricted to records available at admission time $t_0$ and supervised by the diagnoses recorded during the ED encounter rather than at discharge. An LLM auditor further verifies every free-text label as supported, partially supported, or unsupported by that evidence; the primary evaluation scores only fully supported labels. Under a semantic LLM-as-judge protocol, no evaluated system --- frontier general, medical-specialized, or in-domain post-trained --- synthesizes admission-time evidence reliably. Zero-shot models score largely by extraction, recovering only 3-31% of diagnoses that must be inferred rather than read from the record; post-training raises inference-dependent recall to 56%, but a sizeable margin remains, and on time-critical conditions no system attains a clinician's balance of sensitivity and precision. We release the full construction and evaluation pipeline at here.