发表机构
Harvard T.H. Chan School of Public Health; Harvard Medical School; Clalit Health Services; Harvard College; Harvard-MIT Program in Health Sciences and Technology(哈佛陈曾熙公共卫生学院; 哈佛医学院; 克拉利特医疗服务机构; 哈佛学院; 哈佛-麻省理工健康科学与技术项目)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建含208个罕见病临床情景的基准,发现11种先进LLMs在决策中优先资源均等分配而非患者获益,且存在权威框架效应,揭示LLMs决策支持系统会忽略精细伦理考量。
AI 中文摘要
临床决策往往需要优先考虑伦理价值观,如行善、不伤害、尊重患者自主和公正。近期研究已开始评估大语言模型(LLMs)如何做出这类主观、充满价值取向的临床判断。然而,针对罕见病诊疗场景中LLMs决策的评估仍有所欠缺,这类场景中伦理冲突无处不在,且稀缺的先验信息可能会影响LLMs的行为。本研究提供了一个包含208个基于临床实际的罕见病 vignette(临床情景案例)的基准数据集,每个案例都呈现了真实的高风险冲突。当提示11种最先进的LLMs在这些案例中具有临床合理性但存在伦理冲突的后续步骤间做出选择时,我们发现所有被评估的模型始终将公正置于其他核心生物伦理原则之上。具体而言,模型压倒性地倾向于资源均等分配而非基于需求的考量,这表明LLMs对临床严重程度或情境背景的差异反应有限。我们还识别出一种强烈的权威框架效应:在基于委员会的决策场景中,模型倾向于公正;而当最终决策框架为临床医生或患者分别做出时,模型则转向行善和自主。本研究表明,围绕罕见病资源利用的机构压力可能会无声地反映在基于LLMs的决策支持系统中,而更精细的伦理考量则被忽略。
英文摘要
Clinical decision-making often involves prioritizing ethical values, such as beneficence, non-maleficence, respecting a patient's autonomy, and justice. Recent work has begun to assess how large language models (LLMs) make such subjective, value-laden clinical judgments. However, evaluations of LLM decision-making in rare disease care contexts, where ethical tensions are ubiquitous and where scarce prior information likely impacts LLM behavior, are still lacking. Here, we present a benchmark of 208 clinically grounded rare disease vignettes, each of which presents genuine, high-stakes conflicts. When prompting 11 state-of-the-art LLMs to choose between clinically defensible yet ethically conflicting next steps embedded within these vignettes, we found that all evaluated models consistently prioritized justice over other core bioethical principles. Specifically, models overwhelmingly favor equal resource allocation over need-based considerations, indicating LLMs' limited responsiveness to differences in clinical severity or situational context. We also identify a strong authority-framing effect: models favor justice in committee-based contexts and shift toward beneficence and autonomy only when final decisions are framed as being made by clinicians or patients respectively. Our work suggests that institutional pressures surrounding rare disease resource utilization may be silently reflected in LLM-based decision support systems, with finer ethical considerations disregarded.