arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12339cs.CLcs.AIcs.HC

无理解的模仿:大语言模型中决策偏差的起源

Mimicry without understanding: the origins of decision bias in large language models

Eldad Yechiam, Adi Tarabeih

AI总结:

本研究揭示大语言模型(LLMs)存在两类偏差生成机制,经四项经济偏差实验发现ChatGPT-4o和Qwen会表现出社会认同偏差与损失厌恶偏差,偏差程度可被科学报告预测,还明确了LLMs偏差的潜在组成过程。

AI中文摘要:

研究发现大语言模型(LLMs)易受多种社会、情感及认知偏差影响。本文探讨了两种偏差生成机制,即便训练数据中的人类偏好无偏差或被正确归类为偏差,偏差仍可由此产生:第一种是基于人类行为对偏好的错误模仿,即LLMs会在行为与偏好逻辑无关时仍推断人类偏好;第二种是对明确带偏差的人类行为的模仿。在四项聚焦经济偏差的研究中,我们发现ChatGPT-4o与通义千问(Qwen)在被提示提供与个体实际偏好明显无关的人类行为报告时,仍表现出社会认同偏差;当被明确提示损失厌恶为一种偏差时,LLMs也会表现出该偏差。实际上,当被提示提供详细科学报告时,报告中的偏差程度(即损失厌恶)可预测LLMs自身后续的偏差程度。因此,关于偏差的科学论文可成为自我实现的预言,至少在LLMs的响应方面如此。本研究不仅明确了LLMs的偏差,还揭示了其潜在的组成过程。

英文摘要:

Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases. We examined two mechanisms through which such biases can be generated even when human preferences (in the training data) are not biased or when they are correctly categorized as being biased. The first is faulty mimicry of preferences based on human behavior: this involves LLMs inferring human preferences even when behaviors are logically unrelated to preferences. The second is mimicry of explicitly biased human behaviors. In four studies focusing on economic biases, we find that ChatGPT-4o and Qwen exhibited social proof biases even when prompted with reports of human behaviors that were clearly non-indicative of individuals' actual preferences. LLMs also displayed loss aversion when it was explicitly described as a bias. Indeed, when prompted with detailed scientific reports, the extent of the bias (i.e., loss aversion) in the scientific report predicted LLMs' own subsequent bias. Scientific papers of biases can thus become self-fulfilling prophecies, at least when it comes to LLMs' responses. The current study goes beyond fleshing out LLM biases and sheds light on the underlying component processes.

补充信息

↑