合成医院:一个开放、可验证、经医生验证的纵向电子健康记录基准
Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark
浏览论文内容
中文总结 AI 辅助
合成医院是一个开放、可验证、经医生验证的纵向EHR基准,基于公共医学教育材料构建,包含1,268名患者和5,602次就诊,用于测试临床AI性能,结果显示当前模型远未达到上限。
中文摘要 AI 辅助
前沿语言模型很少用于临床工作流程,因为开发它们所需的现实纵向基准十分稀缺。真实的电子健康记录(EHR)数据因隐私、伦理或数据使用问题无法公开共享,且由于病历记录仅反映临床医生的记录内容,其不包含可验证的黄金标准。我们引入了合成医院(Synthetic Hospital),一个开放、完全合成、基于事实的纵向EHR基准,解决了公开共享和可验证黄金标准的障碍。该基准完全基于公共医学教育材料构建,不包含任何受保护的健康信息,包含1,268名纵向患者和5,602次就诊记录,其中每项诊断、发现和时间关系均基于标准本体(ICD-10-CM、SNOMED CT、LOINC),并具有完整的溯源链,可追溯至其来源的医学教育材料。合成医院通过一个模拟医院记录系统提供服务,该系统镜像了真实EHR基础设施(标准互操作性API、基于角色的访问和函数调用接口)。在盲审中,医生以接近随机水平的准确率(53%)区分其记录与真实患者病历。在10个前沿和开放模型中,没有一个接近上限:最佳模型重建患者纵向问题列表的严重性加权F1分数为0.73,与匹配子集上七名医生的平均水平相当,但远低于其中最佳者(0.89),且在总结病历时遗漏了约一半的临床相关发现。总体而言,这些结果凸显了合成医院是对临床AI性能的一个困难且现实的测试。
英文摘要
Frontier language models are rarely used in clinical workflows because the realistic, longitudinal benchmarks needed to develop them are scarce. Real electronic health record (EHR) data cannot be openly shared due to privacy, ethics or data use issues and it does not contain verifiable ground truth since the chart records only reflect what clinicians documented. We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers. Built entirely from public medical-education material with no protected health information, it comprises 1,268 longitudinal patients and 5,602 encounters, where every diagnosis, finding, and temporal relation is grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) and with a complete provenance chain back to its source medical education material. Synthetic Hospital is served through a simulated hospital record system that mirrors real EHR infrastructure (standard interoperability APIs, role-based access and function-calling interface). In a blinded review, physicians distinguished its records from real patient charts at near-chance rates (53\%). Across 10 frontier and open models, none approaches ceiling: the best model reconstructs a patient's longitudinal problem list with a severity-weighted F1 of 0.73, level with the mean of seven physicians on a matched subset but well below the best of them (0.89), and misses roughly half of clinically relevant findings when summarizing a chart. Overall, these results highlight that Synthetic Hospital is a difficult and realistic test of clinical AI performance.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。