arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CLIMB:通过多轮对话诊断共病情况的临床多病种基准

CLIMB: A Clinical Multimorbidity Benchmark for Diagnosing Co-occurring Conditions through Multiturn Conversations

Yusuf Kesmen, Aniruddha Mukherjee, Yena Chang, David Sasu, Trevor Brokowski, Alexandra V. Kulinkina, Kristina Keitel, Akhil Arora, Lars Henning Klein, Mary-Anne Hartley

arXiv 2609.35462首次发表:更新:

发表机构

EPFL; Aarhus University(洛桑联邦理工学院; 奥胡斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CLIMB基准,通过多轮对话模拟患者来评估模型诊断共病的能力,发现现有模型性能低下且表现出单假设追踪行为,难以准确识别多种共病。

AI 中文摘要

患者常常同时患有多种临床疾病,而识别和区分这些疾病所需的发现会在诊疗过程中逐渐显现。在此情境下评估临床推理能力需要多轮交互和多标签诊断。我们引入了CLIMB,一个基准,其中医生模型通过访谈模拟患者来恢复一组真实存在的共病情况。病例由临床决策算法和诊断数据集合成,将多病种表现建立在结构化临床知识之上。在六个前沿和开放模型中,没有一个模型在超过10%的交互案例中准确恢复出全部疾病组合。当疾病共病时,即使模型获得完整的临床记录和真实的疾病数量,诊断性能也会下降。交互进一步降低了性能。在受控实验中,模型表现得像单假设追踪器:它们锚定于开场发现所提示的诊断,围绕该诊断持续提问,并且仅当视野中的发现指向第二种疾病时才恢复该疾病。进一步追问并不能完善疾病集合,反而增加了大部分错误的诊断。我们通过一个单假设追踪的理论参考模型形式化了这一模式。基准、生成器和评估代码可在以下网址获取:此 https URL。

英文摘要

Patients often have several co-occurring clinical conditions, and the findings needed to identify and disambiguate them emerge over the course of a consultation. Evaluating clinical reasoning in this setting requires both multi-turn interaction and multi-label diagnosis. We introduce CLIMB, a benchmark in which a doctor model interviews a simulated patient to recover a ground truth set of co-occurring clinical conditions. Cases are synthesized from clinical decision algorithms and diagnostic datasets, grounding multimorbid presentations in structured clinical knowledge. Across six frontier and open models, none recovers the exact set of conditions in more than 10% of interactive cases. Diagnostic performance declines when conditions co-occur, even when models receive the full clinical record and the true number of conditions. Interaction reduces performance further. In controlled experiments, models behave like single-hypothesis trackers: they anchor on the diagnosis suggested by the opening findings, keep questioning around it, and recover a second condition mainly when a finding in view points to it. Questioning them further does not complete the set but adds mostly wrong diagnoses. We formalise this pattern with a theoretical reference model of single-hypothesis tracking. The benchmark, generator, and evaluation code are available at https://anonymous.4open.science/r/CLIMB-8340.

Comments52 pages (9 main text), 23 figures, 22 tables. Preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑