发表机构
Atman Labs; Oxford University Hospitals; University of Oxford; NIHR Oxford Biomedical Research Centre; Imperial College London; Technical University of Munich; Massachusetts General Hospital; Mass General Brigham; Harvard Medical School; Barts Health NHS Trust; the University of Texas at Austin(阿特曼实验室; 牛津大学医院; 牛津大学; 英国国立卫生研究院牛津生物医学研究中心; 伦敦帝国理工学院; 慕尼黑工业大学; 麻省总医院; 麻省总医院布里格姆医疗系统; 哈佛医学院; 巴特保健国民保健信托基金会; 德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文指出,尽管LLM在医学考试和部分病例推理上表现出色,但因存在信息收集缺陷等问题,尚不适合用于无医生参与的自主临床分诊。
AI 中文摘要
大型语言模型(LLM)目前已能通过医学执照考试,在精心整理的病例中可在诊断推理方面与医生相媲美。这些进展推动了LLM在症状评估、临床决策支持(涵盖诊断与治疗指导、行政文书工作及基于规则的警报增强)中的应用。本文的观点聚焦于其中最具影响力的应用:对自行就诊、诊断未明确的患者进行自主分诊,且临床医生极少或完全不参与该过程。针对这一任务,目前尚缺乏安全性证据。这一差距并非在于医学知识,而在于临床评估的保真度:以生成最可能文本为优化目标的模型,无法在安全答案为低概率的“必须排查的漏诊诊断”时做出安全决策。安全分诊并非选择最可能的诊断,而是在非对称成本下的序列决策,其中单次灾难性漏诊远重于多次误诊,且关键信号可能是患者未主动提供、模型也未被训练去主动探寻的信息。因此,核心缺陷在于不确定性下的信息收集能力不足。在病史不完整的情况下,LLM系统可能无法展现安全分诊所需的行为:拓宽鉴别诊断范围、探寻缺失的危险信号、降低升级处理的阈值、在获取足够信息前暂缓判断,以及在高危害诊断仍未被排除时提升关注程度。考虑到目前的评估常采用完整、精心整理、置信度可控的模拟病例,LLM的这类失败模式可能难以被察觉。若不受临床分诊逻辑约束,LLM在上述场景下的应用可能会因类助手行为和正向偏差(包括轻信、随和及校准不当)而被放大。
英文摘要
LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the use of LLMs for symptom assessment and clinical decision support in diagnostic and treatment guidance, administrative documentation, and rules-based alert enhancement. This Perspective concerns the most consequential of these applications: the autonomous triage of self-presenting, undifferentiated patients, with little or no clinician in the loop. For that task, the evidence of safety does not yet exist. The gap is not in medical knowledge but in the fidelity of clinical evaluation: a model optimized to continue the most probable text is not optimized to act safely when the safe answer is the improbable must-not-miss diagnosis. Safe triage is not the selection of the most likely diagnosis; it is a sequential decision under asymmetric cost, in which the single catastrophic miss outweighs many false alarms, and the decisive signal may be one the patient has not volunteered - and that the model has not been trained to seek. The core deficit is therefore one of information gathering under uncertainty. Under incomplete histories, LLM systems may fail to show the behaviors safe triage requires: broadening the differential; seeking the missing red flag; lowering the threshold for escalation; deferring judgement until sufficient information is obtained; and escalating concern where high-harm diagnoses remain unexcluded. These modes of failure for LLMs can be difficult to detect considering that evaluations to date often use complete, well-curated, confidence-gated simulations. The application of LLMs under these conditions may be amplified by assistant-like behaviors and positive bias, including credulity, agreeableness, and miscalibration - when these are not constrained by clinical triage logic.
Comments3 figures