发表机构
University of Illinois Urbana-Champaign; NYC Health + Hospitals/Jacobi Medical Center; University of Chicago(伊利诺伊大学厄巴纳-香槟分校; 纽约市健康与医院公司雅各比医学中心; 芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究开发决策理论框架研究LLM的修辞失配问题,通过临床实验发现其会诱导平均2.81%的有害决策翻转,还提出用LLM模拟决策者实现可扩展评估,揭示了高风险领域的新安全隐患。
AI 中文摘要
人类决策常受诸多已被充分记录的认知偏差影响。随着大语言模型(LLMs)越来越多地融入高风险人机决策场景,重要的是要了解其输出是否会放大潜在偏差、这如何影响人类决策,以及关键的是,是否会导致有害后果。在这项工作中,我们开发了一个决策理论框架来研究修辞失配,这是一种故障模式,即LLM在给定决策情境下使用修辞上不恰当的呈现形式,从而诱导人类做出次优决策。我们通过一项使用美国医师执照考试整理的数据集开展的现实临床决策人类受试者实验,对这一现象进行实证研究。通过测量LLM生成的信息如何影响决策,我们观察到,在不同模型中,LLM诱导出平均2.81%的有害决策翻转率,即临床医生参与者从正确答案改为错误答案。参与者报告的理由提供了证据,表明这些修改与LLM使用的语言密切相关,可能会诱发不同类型的认知偏差,包括锚定效应、权威偏差和损失厌恶。为实现可扩展评估,我们使用LLM模拟的决策者实例化我们的理论框架,以计算衡量修辞失配。我们的研究结果揭示了高风险领域中此前未被认识到的安全问题:模型可能在事实上对齐,但仍会通过其修辞呈现方式造成伤害。
英文摘要
Human decision-making is often shaped by a range of well-documented cognitive biases. As large language models (LLMs) become increasingly integrated into high-stakes human-AI decision-making, it is important to understand whether their outputs can amplify potential biases, how this influences human decisions, and crucially, whether it can lead to harmful consequences. In this work, we develop a decision-theoretic framework to study rhetorical misalignment, a failure mode where an LLM uses rhetorically inappropriate forms of presentation for a given decision context, thereby inducing suboptimal human decisions. We empirically investigate this phenomenon through a human-subject experiment in realistic clinical decision-making using a dataset curated from the United States Medical Licensing Examination. By measuring how LLM-generated information affects decisions, we observe that LLMs induce an average 2.81% rate of harmful decision flips across different models, where clinician participants change from a correct to an incorrect answer. Rationales reported by participants provide evidence that these revisions are closely related to the language used by LLMs that may induce different types of cognitive biases, including anchoring, authority bias, and loss aversion. To enable scalable evaluation, we instantiate our theoretical framework using decision-makers simulated by LLMs to computationally measure rhetorical misalignment. Our findings reveal a safety concern previously unrecognized in high-stakes domains: a model can be factually aligned yet still induce harm through its rhetorical presentation.