发表机构
Zhejiang University; Alibaba Group; Transvascular Implantation Devices Research Institute(浙江大学; 阿里巴巴集团; 血管内植入装置研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
介绍基于EHR的LongMedBench基准,用于长期临床决策。构建含多患者多事件数据集,提出评估分类法。实验表明大语言模型在隐式时间推理有挑战,RAG和智能体记忆系统对信息检索有帮助,决策任务性能依赖模型即时上下文。
AI 中文摘要
在这项工作中,我们引入了LongMedBench,这是一个基于真实世界电子健康记录(EHR)的长期临床决策基准。此前对基于大语言模型(LLM)的医疗智能体的评估主要强调短上下文知识问答和工具使用。但现实医疗本质上是纵向的,临床医生必须汇总多次就诊、检查和不断变化的治疗中的证据。因此,长期交互对于现实评估至关重要。LongMedBench通过可重复的管道构建,将MIMIC-IV入院记录和临床笔记整合到时间序列事件流和长上下文记忆数据集中,实现智能体与临床环境之间的长期、多会话交互。它包含335名患者,平均每名患者有19.72次住院就诊,每次就诊有44.91个医疗事件。在长期决策过程的指导下,我们提出了一个包含三个套件的评估分类法:基于事实的问答、时间推理和长期决策。该分类法衡量智能体如何在更长时间范围内理解和利用历史患者信息。我们的实验表明,虽然最近的大语言模型可以很好地利用显式时间戳,但在隐式时间推理方面存在挑战;检索增强生成(RAG)和智能体记忆系统可以提高信息检索任务的性能,但决策任务的性能高度依赖于模型的即时上下文。
英文摘要
In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making. Prior evaluations of LLM-based medical agents have largely emphasized short-context knowledge QA and tool use. However, real-world medical care is inherently longitudinal, and clinicians must aggregate evidence across repeated visits, tests, and evolving treatments. Therefore, long-horizon interaction is essential for realistic assessment. LongMedBench is constructed via a reproducible pipeline that integrates MIMIC-IV admission records and clinical notes into time-series event streams and long-context memory datasets, enabling long-horizon, multi-session interactions between agents and a clinical environment. It comprises 335 patients, with 19.72 inpatient visits per patient on average and 44.91 medical events per visit. Guided by the long-horizon decision process, we propose an evaluation taxonomy with three suites: fact-based QA, temporal reasoning, and long-horizon decision-making. This taxonomy measures how agents understand and leverage historical patient information over extended horizons. Our experiments show that while recent LLMs can make good use of explicit timestamps, they have challenges in implicit time inference; The RAG and agent memory system can improve the performance of information retrieval tasks, but the performance of decision-making tasks is highly dependent on the model's immediate context.
CommentsSubmitted manuscript prior to peer review in MICCAI 2026