arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20695cs.CY

语言模型体现并放大人类认知扭曲:该怎么办?

Language Models Embody and Amplify Human Cognitive Distortions: What Is to Be Done?

Arnau Marin-Llobet, Steven A. Lehr, Mahzarin R. Banaji

AI总结:

研究指出语言模型存在社会认知偏见,它不仅体现人类偏见,还会放大、加剧并传递偏见,鉴于其对决策影响巨大,提出了从诊断、监管和操作方面应对语言模型偏见问题的对策。

AI中文摘要:

人类判断从根本上容易出错。人工智能的一个承诺是它将消除决策中的偏见,为所有人确保一个更公平、更安全的世界。然而研究明确表明,语言模型表现出重大的社会认知偏见。我们提醒读者,人工智能中的偏见:(a)是隐蔽的,具有讽刺意味的是它是对齐目标的一个特征;(b)不仅是一面镜子,更是人类偏见的放大器;(c)在模型迭代中加剧;(d)甚至将偏见传递给人类。鉴于人工智能对决策可能产生的巨大且普遍的影响,我们提出了诊断、监管和操作方面的对策。

英文摘要:

Human judgment is fundamentally prone to error. A promise of AI is that it will rid decisions of bias and ensure a fairer and safer world for all. Yet research unequivocally demonstrates that LLMs exhibit consequential sociocognitive biases. We alert readers that bias in AI (a) is covert and ironically a feature of alignment goals, (b) is not merely a mirror, but an amplifier of human bias, (c) intensifies across model generations, and (d) even transmits bias to humans. Given the potentially seismic and ubiquitous influence of AI on decision making, we propose countermeasures that are diagnostic, regulatory and operational.

↑