发表机构
University of Texas at Dallas(德克萨斯大学达拉斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对临床LLM在不完整证据下的决策,提出DECIDE/ASK/DEFER三行动策略,通过盲法协议和DDXPlus数据集评估,发现其能揭示二元弃权隐藏的行为,但未产生统一更优的决策策略。
AI 中文摘要
临床大语言模型不仅需要决定产生何种诊断,还必须判断现有证据是否足以支持自主决策。二元的DECIDE/ABSTAIN(决定/弃权)表述将不同的非决策状态合并,并未明确评估信息获取。我们引入了一种DECIDE/ASK/DEFER(决定/询问/转交)的表述,并配以盲法协议,防止模型利用证据完整性元数据。我们在由DDXPlus构建的200个匹配的临床证据状态上评估了Qwen、Gemini和GPT。模型在相同证据下的行动选择表现出显著差异,在200个状态中有137个存在分歧。对于Qwen,匹配的目标性分析与随机分析表明,所选信息改变后续自主决策可能性的程度比诊断正确性更明显。其匹配的DECIDE/ABSTAIN基线进一步揭示了安全性与自主性之间的权衡:三行动策略挽救了一些错误的自主决策,但也移除了一些正确的自主决策。这些结果表明,将信息获取与临床医生转交分离,能够暴露二元弃权所隐藏的行为,但并未产生统一改进的决策策略。
英文摘要
Clinical LLMs must decide not only what diagnosis to produce, but also whether the available evidence is sufficient for autonomous decision making. Binary DECIDE/ABSTAIN formulations merge distinct non decision states and do not explicitly evaluate information acquisition. We introduce a DECIDE/ASK/DEFER formulation together with a blinded protocol that prevents models from using evidence completeness metadata. We evaluate Qwen, Gemini, and GPT on 200 matched clinical evidence states constructed from DDXPlus. The models show substantial differences in action selection under identical evidence, with disagreement in 137 of 200 states. For Qwen, a matched targeted versus random analysis shows that selected information changes the likelihood of a subsequent autonomous decision more clearly than diagnostic correctness. Its matched DECIDE/ABSTAIN baseline further reveals a safety autonomy tradeoff: the three action policy rescues some erroneous autonomous decisions but also removes some correct autonomous deci sions. These results show that separating information acquisition from clinician deferral exposes behavior that binary abstention hides, without yielding a uniformly improved decision policy.
CommentsAccepted to the Main Track of the NeurIPS 2026 GenAI4Health Workshop