发表机构
University of Massachusetts Lowell; University of Massachusetts Amherst(马萨诸塞大学洛厄尔分校; 马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对出院教育中LLM评估忽视患者理解的问题,提出DischargeBench人物设定模拟框架,通过虚拟患者多轮对话和四维评分,揭示总分掩盖的临床差异,强调评估应聚焦患者理解。
AI 中文摘要
出院教育是一项互动式教学任务:临床医生需根据患者的识字能力、记忆力和个性调整出院计划。现有的LLM评估主要针对静态或文档生成任务,并未衡量开放式对话中患者的理解程度。我们引入了DischargeBench,一个基于人物设定的模拟框架,其中候选LLM教育者与虚拟患者进行多轮对话,同时一个教育监控代理在不修改教育者的情况下调节患者的真实性,从而保护评估信号。我们整理了MIMIC-IV-Ext-DischargeBench数据集,涵盖24个ICD章节的477个病例,并包含人物设定轴(个性、教育水平、健康素养、既往病史回忆)以进行分层分析。每次模拟由LLM作为裁判,依据与医生注释对齐的标准,从四个维度评分:对话质量、主题清单、理解度和事实一致性。在闭源和开源LLM中,总分掩盖了不同ICD章节和患者人物设定间具有临床相关性的差异;困难的人物设定暴露了覆盖不足、理解差距以及源答案一致性降低的问题。出院教育的LLM评估应聚焦于患者理解,而非仅关注文本质量或答案准确性。
英文摘要
Hospital discharge education is an interactive teaching task: a clinician adapts a discharge plan to a patient's literacy, recall, and personality. Existing LLM evaluations target static or artifact-generation tasks and do not measure patient understanding under open-ended dialogue. We introduce DischargeBench, a persona-grounded simulation in which a candidate LLM educator conducts a multi-turn session with a Virtual Patient, while an Education Monitor Agent regulates patient realism without modifying the educator, protecting the evaluation signal. We curate MIMIC-IV-Ext-DischargeBench, 477 cases over 24 ICD chapters with persona axes (personality, education level, health literacy, past-medical-history recall) for stratified analysis. Each simulation is scored on four axes -- Conversation Quality, Topic Checklist, Comprehension, and Factual Consistency -- by an LLM-as-a-Judge aligned against physician annotations. Across closed- and open-source LLMs, aggregate scores conceal clinically relevant variation across ICD chapters and patient personas; difficult personas expose coverage failures, comprehension gaps, and reduced source-answer agreement. LLM evaluation for discharge education should center patient understanding, not text quality or answer accuracy alone.