arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.15735cs.CLcs.AI

EHRNote-ChatQA:一个面向纵向出院总结的基于证据的多轮临床问答基准

EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering over Longitudinal Discharge Summaries

  • KAIST(韩国科学技术院)
  • Seoul National University(首尔大学)
  • Seoul National University Bundang Hospital(首尔大学盆唐医院)
  • SAIHST, Sungkyunkwan University(成均馆大学)
  • Yonsei University College of Medicine(延世大学医学院)
  • Gangnam Severance Hospital(江南塞弗伦斯医院)
  • Severance Hospital(塞弗伦斯医院)
  • Seoul Medical Center(首尔医疗中心)
  • Seoul National University Hospital(首尔大学医院)
  • National Cancer Center(国立癌症中心)
  • Icahn School of Medicine at Mount Sinai(西奈山伊坎医学院)
  • Samsung Medical Center(三星医疗中心)

机构由 AI 辅助整理,请以论文原文为准。

Jiyoun Kim, Muhan Yeo, Eunhye Jang, Jeewon Yang, Hangyul Yoon, Su Ji Lee, Hee Jo Han, Hee-Jae Jung, Doyun Kwon, Jun young Lee, Jaehun Lee, Jung-Oh Lee, Sunjun K… 展开作者

Jiyoun Kim, Muhan Yeo, Eunhye Jang, Jeewon Yang, Hangyul Yoon, Su Ji Lee, Hee Jo Han, Hee-Jae Jung, Doyun Kwon, Jun young Lee, Jaehun Lee, Jung-Oh Lee, Sunjun Kweon, Jong Hak Moon, Daseul Kim, Minjae Cho, Edward Choi

AI总结:

提出EHRNote-ChatQA基准,基于MIMIC-IV出院总结构建,包含967个多轮样本和16072个专家验证的QA对,评估LLM在证据支持下的多轮临床问答能力,发现模型在证据定位和多轮错误累积方面存在挑战。

AI中文摘要:

出院总结是关键的临床文档,包含患者整个住院期间的背景信息,医疗专家在患者再入院、持续护理和诊断决策中会常规审阅这些文档。在审阅时,医疗专家通常必须迭代地综合多个总结中的信息,同时验证支持每个答案的证据。尽管大型语言模型(LLM)在临床问答中的应用日益增多,但现有基准未能充分反映这一场景:它们通常评估考试式的医学知识,或侧重于单轮问答且证据定位评估有限。我们引入了EHRNote-ChatQA,这是首个针对患者多个出院总结的基于证据的多轮临床问答基准。该基准基于去标识化的MIMIC-IV出院总结构建,包含967个患者级多轮样本,涵盖1到5份笔记,以及16072个经医学专家验证的QA对(8036个内容问题,每个配对有一个证据定位问题),覆盖八个临床类别。基准通过专家指导的流程构建,结合出院总结结构化模式、专家策划的多轮QA模板和基于LLM的生成,随后由11位医学专家对每个QA样本进行审查和修订。对22个开源和闭源LLM的基准测试揭示了若干挑战,包括LLM在证据定位方面比内容回答更困难、多轮错误随轮次累积,以及单轮临床QA性能无法可靠迁移到该场景。这些发现确立了EHRNote-ChatQA作为评估临床QA系统的严格且实用的基准。该数据集将通过PhysioNet凭证访问公开发布。

英文摘要:

Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision-making. When reviewing them, medical experts often must iteratively synthesize information across multiple summaries while verifying the evidence supporting each answer. Although large language models (LLMs) are increasingly explored for clinical question answering, existing benchmarks do not sufficiently reflect this setting: they often evaluate exam-style medical knowledge or focus on single-turn question answering with limited evidence-grounding evaluation. We introduce EHRNote-ChatQA, the first benchmark for evidence-grounded multi-turn clinical question answering over patients' multiple discharge summaries. Built from de-identified MIMIC-IV discharge summaries, EHRNote-ChatQA contains 967 patient-level multi-turn samples spanning one to five notes and 16,072 medical-expert-verified QA pairs (8,036 content questions, each paired with an evidence-grounding question) across eight clinical categories. The benchmark is constructed through an expert-informed pipeline combining discharge-summary structuring schema, expert-curated multi-turn QA templates, and LLM-based generation, followed by review and revision of every single QA sample by 11 medical experts. Benchmarking 22 open- and closed-source LLMs reveals several challenges, including that LLMs struggle more with evidence grounding than content answering, multi-turn errors compound across turns, and single-turn clinical QA performance does not reliably transfer to this setting. These findings establish EHRNote-ChatQA as a rigorous and practical benchmark for evaluating clinical QA systems. The dataset will be made publicly available through PhysioNet credentialed access.

↑