arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10275cs.AIcs.CL

大语言模型在代理临床推理中的信息获取失败

Information-seeking failures of large language models in agentic clinical reasoning

Krischan Braitsch, Laura K. Schmalbrock, Theresa Weltermann, Andrew F. Berdel, Isabella Miller, Kai Tran, Michael Heider, Sabrina Kraus, Florian Bassermann, Jac… 展开作者

Krischan Braitsch, Laura K. Schmalbrock, Theresa Weltermann, Andrew F. Berdel, Isabella Miller, Kai Tran, Michael Heider, Sabrina Kraus, Florian Bassermann, Jacqueline Lammert, Sebastian Ziegelmayer, Marcus Makowski, Lisa C. Adams, Keno K. Bressem

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型在临床推理中信息获取失败问题,开发血液肿瘤学代理评估框架,发现信息利用率是诊断准确性关键指标,推理痕迹与准确性不相关,主要失败模式为搜索满足等,指出模型主要限制是不确定性下信息获取失败。

中文摘要 AI 辅助

大语言模型在医学知识评估中取得高分,但临床推理需要在不确定性下积极决定调查内容。我们在血液肿瘤学中开发了一个代理评估框架,模型在做出诊断和治疗计划前必须在三个连续轮次中主动请求临床数据。在32个前沿模型中,最佳模型总体准确率仅为68%。信息利用率是诊断准确性的最强预测指标,但在最后一轮从57%降至26%,导致治疗选择关键的分子和细胞遗传学数据未被审查。推理痕迹在临床推理评分标准上得分高,但与准确性不相关。错误分析确定搜索满足、锚定和过早关闭是主要失败模式,与新手临床医生在诊断推理双过程模型下的认知偏差相同。这些发现表明当前临床肿瘤学模型的主要限制不是医学知识不足,而是在不确定性下信息获取的系统性失败。

英文摘要

Large language models achieve high scores on medical knowledge assessments, yet clinical reasoning requires actively deciding what to investigate under uncertainty. We developed an agentic evaluation framework in hematologic oncology in which models must proactively request clinical data across three sequential rounds before committing to a diagnosis and treatment plan. Across 32 frontier models, the best achieved only 68% overall accuracy. Information utilization, the fraction of available data actually requested, was the strongest predictor of diagnostic accuracy (R = 0.69, P < 0.001), yet utilization collapsed from 57% to 26% in the final round, leaving molecular and cytogenetic data critical for treatment selection unexamined. Reasoning traces scored high on a clinical reasoning rubric (91% above threshold) but decorrelated from accuracy, revealing a gap between locally coherent rationales and globally correct conclusions. Error analysis identified search satisficing, anchoring and premature closure as the dominant failure modes, the same cognitive biases that characterize novice clinicians under dual-process models of diagnostic reasoning. These findings demonstrate that the primary limitation of current models in clinical oncology is not insufficient medical knowledge but a systematic failure of information-seeking under uncertainty.

补充信息

↑