发表机构
Massachusetts Institute of Technology; Primitive Labs(麻省理工学院; Primitive实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过对OpenAI和Claude模型的实验,发现LLM智能体的群体偏袒源于观察到的群体行为,利害关系下会被个体声誉凌驾,还能预测其对背叛等场景的反应。
AI 中文摘要
语言模型智能体偏向自身群体,是因为它们观察到群体成员间会互相偏袒。一旦决策涉及成本,仅群体标签的作用微乎其微;驱动偏袒的是观察到的行为,而个体自身的记录可凌驾于群体标签之上。我们在带有任意群体标签的小型社会中对此展开测试,设置十轮点数分享,对15个OpenAI模型和3个Claude模型进行匹配的单次决策,共涉及约4400个社会和330万次经审计的模型调用。首先,早期研究中提到的纯群体标签的显著效应,仅在给予他人点数对智能体无成本时才会出现;一旦智能体可自行保留点数,所有表现出该效应的模型的这一效应均会消失。其次,在有利害关系的情况下,互动历史成为偏袒的主要来源:15个模型中有13个的历史效应具有统计显著性,其中11个模型的历史效应达到10分中的3.5至8分,且随互动轮次增加而增强,还会延伸至智能体从未接触过的带标签陌生人。第三,在预设历史下,平等主义规范下偏袒降至接近零,而当观察到自身群体偏向另一方时,偏袒会反转;当个体记录与群体冲突时,更强的模型会站在个体记录一侧。因此,群体偏袒是对观察到的群体行为的遵从,通过标签传递给陌生人,且可被个体声誉凌驾。该解释同样可预测对背叛、丑闻和免费换群体的反应:公开谴责修复背叛的效果优于道歉或赔偿,分配惩罚仅针对违规成员,已形成的群体无法被收买,但较弱模型可能会被邀请脱离原群体。
英文摘要
Language-model agents favor their own group because they have watched their members favor each other. The group label alone does little once the decision has a cost; what drives favoritism is observed behavior, and an individual's own record can override it. We test this in small societies with arbitrary group labels, ten rounds of point sharing, and matched one-shot decisions across fifteen OpenAI models and three Claude models, about 4,400 societies and 3.3 million audited model calls. First, the large effect of a bare group label reported in earlier work appears only when giving others points costs the agent nothing; once the agent can keep points for itself, that effect collapses on every model that shows it. Second, under a stake, interaction history becomes the main source of favoritism: the history effect is statistically positive on 13 of 15 models, reaches about 3.5-8 points out of 10 on 11, grows with the number of rounds played, and extends to labeled strangers the agent has never met. Third, with scripted histories, favoritism falls to near zero under an egalitarian norm and reverses when the agent's own group is seen favoring the other side; stronger models side with an individual's record when it conflicts with the group. Group favoritism is thus conformity to observed group behavior, carried to strangers by the label and overridden by individual reputation. The same account predicts responses to betrayal, scandal, and a free offer to change group: public reprimand repairs betrayal better than apology or restitution, allocation punishment remains confined to the offending member, and a formed group cannot be bought but can, on weaker models, be invited away.