arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于语言模型的高风险决策:来自急诊分诊的见解

High-Stakes Decisions with Language Models: Insights from Emergency Triage

Khurram Yamin, Christopher Kelly, Bryan Wilder, Eric Horvitz

arXiv 2608.01361首次发表:更新:

发表机构

Microsoft; Carnegie Mellon University(微软; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究以急诊分诊为案例,将语言模型的高风险决策置于概率决策框架内,发现语言模型可根据效用调整建议,强调高风险场景下需将其作为概率决策系统评估。

AI 中文摘要

不确定性下的高风险决策,例如医疗急诊分诊,仅靠准确预测是不够的,还需要估计替代结果的可能性,同时明确权衡不同行动的后果,这些原则长期以来一直是医学诊断和决策的基础。然而,语言模型越来越多地被用于高风险临床建议,却没有明确规定支配这些决策的效用函数。在此,我们表明,使用语言模型进行急诊分诊可以在概率决策框架内得到理解,为更广泛的决策分析范式提供了一个案例研究,该范式用于在高风险环境中引导、评估和部署语言模型。我们使用来自消费者分诊系统结构化评估的临床 vignette(临床案例),分析了在指定漏诊急诊与不必要升级的相对成本的替代效用函数下的治疗建议。我们发现,能力较强的语言模型会根据规定的效用调整建议,这表明相同的基础预测可以支持截然不同的决策策略。这些发现表明,有效部署不仅取决于改进预测,还取决于明确决策目标。更广泛地说,它们表明,用于高风险应用的语言模型应被理解和评估为概率决策系统,其建议同时取决于预测性能和明确的效用函数。

英文摘要

High-stakes decisions under uncertainty, such as medical emergency triage, require more than accurate predictions. They depend on estimating the likelihood of alternative outcomes while explicitly weighing the consequences of different actions, principles that have long formed the foundation of medical diagnosis and decision making. Yet language models are increasingly used for high-stakes clinical recommendations without explicit specification of the utilities governing these decisions. Here we show that emergency triage with language models can be understood within a probabilistic decision framework, providing a case study of a broader decision-analytic paradigm for steering, evaluating, and deploying language models in high-stakes settings. Using clinical vignettes from a structured evaluation of a consumer triage system, we analyze recommendations for treatment under alternative utility functions that specify the relative costs of missed emergencies and unnecessary escalation. We find that capable language models adjust recommendations in response to stated utilities, revealing that the same underlying predictions can support markedly different decision policies. These findings show that effective deployment depends not only on improving predictions but also on making decision objectives explicit. More broadly, they suggest that language models for high-stakes applications should be understood and evaluated as probabilistic decision systems whose recommendations depend jointly on predictive performance and explicit utilities.

Comments29 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑