arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MedLoCoMo:用于大语言模型的长上下文多会话医学对话基准

MedLoCoMo: A Long-Context Multi-Session Medical Dialogue Benchmark for Large Language Models

Zeyu Zhang, Ziqing Wang, Kaize Ding

arXiv 2607.22566首次发表:更新:

发表机构

Northwestern University(西北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型在多入院医学对话临床推理方面的不足,构建MedLoCoMo基准,通过特定方法生成问答项目,含100个患者时间线,发现跨入院推理比局部证据使用难,提供代码和基准供使用与复现。

AI 中文摘要

MedLoCoMo是一个用于多入院医学对话中特定患者临床推理的医学长上下文记忆基准。现有医学问答基准大多测试短上下文知识或单文档基础,大语言模型能否使用、连接和舍弃纵向患者病史尚待研究。我们通过构建入院级临床数据包、合成有基础的医患对话,并在单入院、跨入院和对抗性无法回答的设置中生成有证据关联的问答项目,从去识别化的MIMIC-IV和MIMIC-IV-Note记录构建MedLoCoMo。该基准包含100个患者时间线,平均每次对话1669.8轮、29.7个会话和74512.2个 tokens。跨入院推理始终比局部证据使用更难。代码和基准可在指定网址获取。

英文摘要

MedLoCoMo is a Medical Long-Context Memory benchmark for patient-specific clinical reasoning over multi-admission medical dialogue. Existing medical QA benchmarks largely test short context knowledge or single document grounding, leaving open whether LLMs can use, connect, and abstain over longitudinal patient histories. We build MedLoCoMo from deidentified MIMIC-IV and MIMIC-IV-Note records by constructing admission-level clinical packets, synthesizing grounded doctor-patient conversations, and generating evidence linked QA items over single-admission, cross-admission, and adversarial unanswerable settings. The benchmark contains 100 patient timelines averaging 1,669.8 turns, 29.7 sessions, and 74,512.2 tokens per conversation. Across the evaluated baselines, cross-admission reasoning is consistently harder than localized evidence use, even when models have long context windows or use external memory or retrieval methods. The code and MedLoCoMo benchmark release is available at https://github.com/leozzy13/MedLoCoMo for use and reproducibility.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑