arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05583cs.CYcs.AIcs.LG

判断-结果差距:医疗决策中的大语言模型道德推理

The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions

Hadi Hosseini, Samarth Khanna, Leona Pierce

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现LLMs在医疗决策中存在判断-结果差距,即虽认同患者对危害健康行为负有责任,但拒绝将该判断用于稀缺医疗资源分配,且随推理能力提升会放大与人类的道德分歧。

中文摘要 AI 辅助

随着大语言模型(LLMs)进入医疗等高风险领域,理解其道德推理变得至关重要。稀缺医疗资源的分配决策往往取决于责任判断,尤其是当患者自身行为导致疾病时。我们研究LLMs如何推理责任及其后果,追踪它们在连续层面的判断:从行为到引发的疾病,再到医疗资源的拒绝分配。我们评估了涵盖不同模型家族和能力水平的多种LLMs,这些模型基于从先前研究改编的各类临床 vignette(临床案例)。我们的研究结果发现了判断-结果差距:LLMs在很大程度上与人类一致,认为患者对危害健康的行为负有责任,但绝大多数情况下拒绝让这一判断影响稀缺医疗资源的分配。具体而言,LLMs默认采用随机分配,而人类始终倾向于选择过错较轻的患者。与人类相比,LLMs也更强调信息获取的重要性,当缺乏健康风险知识时会降低责任判断。这些发现表明,当责任与资源稀缺性交织时,LLMs采用的道德框架与人类存在系统性差异,且令人惊讶的是,随着推理能力提升,LLMs往往会放大与人类的规范性分歧。

英文摘要

As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. Decisions about scarce medical resources often hinge on judgments of responsibility, particularly when patients' own actions contribute to illness. We investigate how LLMs reason about responsibility and its consequences, tracing their judgments across successive levels, from the behavior, to the resulting illness, to the denial of care. We evaluate a wide range of LLMs, spanning different model families and capability levels, on various clinical vignettes adapted from prior studies. Our results identify a judgment-consequence gap: LLMs largely agree with humans that patients bear responsibility for health-harming behaviors, yet overwhelmingly refuse to let that judgment influence how they allocate scarce resources. Specifically, LLMs default to random allocation, whereas humans consistently favor the less-culpable patient. Compared to humans, LLMs also place greater emphasis on access to information, reducing responsibility judgments when health-risk knowledge is unavailable. These findings reveal that LLMs apply a systematically different moral framework than humans when responsibility and resource scarcity intersect, surprisingly often amplifying normative disagreement with humans as reasoning capability increases.

发表机构

  • Pennsylvania State University(宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑