迈向AI的社会心理学:语言模型智能体再现类人类的最小群体偏见
Toward a social psychology of AI: language-model agents reproduce human-like minimal-group bias
AI总结:
研究将社会心理学的最小群体范式改编为探测工具,发现语言模型智能体再现类人类的最小群体偏见,为测量管控AI社会行为提供了新范式。
AI中文摘要:
如今语言模型智能体已能在群体中互动,但仅探测记忆刻板印象内容或用模型模拟人类的评估方式,并未对这种社会行为进行测量。我们将社会心理学经典的群体间偏见测试范式——最小群体范式,改编为受控探测工具:智能体仅依据任意群体标签,在匿名同伴间分配点数。在四个推理模型中,仅被归类为无意义群体就引发了内群体偏爱,这种偏好在无群体的对照组中消失,且集中在数值少数群体:少数群体决策者相对于其人数,对自身群体的分配过多;多数群体决策者的分配接近比例;当群体规模相等时,这种不对称性消失。禁用其中一个模型的推理能力并未消除这种倾向,甚至反而增强了,但几乎抹去了少数群体与多数群体间的不对称性,这表明审议过程影响了偏见的集中位置,而非偏见是否出现。这些开源推理模型再现了人类群体间歧视的行为特征,且独立于刻板印象内容,社会心理学的理论与方法为测量和管控AI的社会行为提供了范式。
英文摘要:
Language-model agents now interact in groups, but evaluations that probe memorised stereotype content or use models to simulate people leave this social behaviour unmeasured. We adapt the minimal-group paradigm---social psychology's classic test of intergroup bias---into a controlled probe: an agent distributes points among anonymous peers bearing only an arbitrary group label. Across four reasoning models, mere categorisation into meaningless groups elicited in-group favouritism that vanished under a group-blind control and was concentrated in the numerical minority: minority deciders over-allocated to their own group relative to their numbers, majority deciders allocated close to proportionally, and the asymmetry closed at equal group sizes. Disabling reasoning in one model did not remove the disposition---if anything it grew---but nearly erased the minority-majority asymmetry, implicating deliberation in where bias concentrates rather than whether it appears. These open-weight reasoning models reproduce the behavioural signature of human intergroup discrimination, independent of stereotype content, and social psychology's theories and methods offer a paradigm for measuring and governing AI's social behaviour.