arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通往同一答案的漫漫长路:大型语言模型在逐步升级的推理预算下的认知偏差

The Long Road to the Same Answer: Cognitive Bias Under Escalating Reasoning Budgets in Large Language Models

Obada Kraishan

arXiv 2610.10049首次发表:更新:

发表机构

College of Media and Communication Texas Tech University(德州理工大学媒体与传播学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过剂量反应实验发现,大型语言模型的推理模型在增加推理预算时并未减少认知偏差,反而可能加剧,表明测试时推理不能保证理性,需逐偏差审计。

AI 中文摘要

推理模型在推理时分配额外的计算资源,并将其答案呈现为深思熟虑的产物。如果这种深思熟虑符合人类认知的双过程理论,那么更长的思考应该会削弱快速、直觉性判断所产生的经典决策偏差。我们使用来自一个成熟基准的30个场景,涵盖六种偏差(锚定、框架、损失厌恶、承诺升级、可得性、确认),在四个模型家族中进行了剂量反应研究,将每个推理模型与匹配的非推理孪生模型配对,并请求思考上限为0、1,024、4,096和8,192个token,共进行了12,350次API调用。由于请求的上限并不等同于实际实现的深思熟虑,我们使用每次调用消耗的推理token作为剂量。首先,推理模型并不比其孪生模型偏差更小;在每个家族中,点估计都倾向于相反方向,但项目级合并对比并不可靠(Delta = +0.031, t(29) = 1.45, p = .157)。其次,偏差幅度并不随实际深思熟虑的增长而可靠下降:没有斜率显著为负,而在有变化的地方,是符号化得分进一步偏离人类方向。第三,锚定是唯一朝向人类方向的偏差(d = 1.89)。其他五种偏差中有四种在所有七个模型中都倾向于相反方向;每个偏差有五个项目,这种反转对框架是可靠的,对承诺升级、确认和损失厌恶是方向性的,而可得性则缺失。在回答前加上一句重新陈述锚定的一行指令,降低了所有五个锚定项目上的锚定效应,而额外的思考无论多少都无法做到这一点,尽管该效应未达到显著性(p = .057)。这些结果反对将测试时推理视为理性的保证,并主张对部署的模型逐偏差进行审计。

英文摘要

Reasoning models allocate extra computation at inference time and present their answers as the product of deliberate thought. If this deliberation works the way dual-process accounts of human cognition suggest, longer thinking should weaken the classic decision biases that fast, intuitive judgment produces. Using 30 vignettes covering six biases (anchoring, framing, loss aversion, escalation of commitment, availability, confirmation) from an established benchmark, we run a dose-response study across four model families, pairing each reasoning model with a matched non-reasoning sibling and requesting thinking ceilings of 0, 1,024, 4,096, and 8,192 tokens, for 12,350 API calls. Because a requested ceiling is not the same as realized deliberation, we use the reasoning tokens each call consumed as the dose. First, reasoning models are not less biased than their siblings; the point estimate leans the other way in every family, but the item-level pooled contrast is not reliable (Delta = +0.031, t(29) = 1.45, p = .157). Second, bias magnitude does not reliably fall as realized deliberation grows: no slope is significantly negative, and where anything moves it is the signed score drifting further from the human direction. Third, anchoring is the only bias in the human direction (d = 1.89). Four of the other five lean the opposite way in all seven models; with five items per bias, that reversal is reliable for framing and directional for escalation of commitment, confirmation, and loss aversion, while availability is absent. A one-line instruction to restate the anchor before answering lowered anchoring on all five anchoring items, which no amount of additional thinking did, although the effect does not reach significance (p = .057). The results argue against treating test-time reasoning as a rationality guarantee and for auditing deployed models bias by bias.

CommentsAccepted at the 2026 IEEE 8th International Conference on Cognitive Machine Intelligence (IEEE CogMI 2026). 8 pages, 3 figures, 5 tables. Code and data: https://github.com/obadaKraishan/anchored-minds

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑