发表机构
Tel-Aviv University; New York University(特拉维夫大学; 纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨多智能体系统中LLMs的从众行为受群体身份影响,发现内群体共识增加从众、外群体共识减少从众,且思维链推理可抑制多数效应,揭示身份因素在AI安全中的操纵风险。
AI 中文摘要
大型语言模型(LLMs)越来越多地被部署在多智能体环境中,其中智能体相互观察和影响,使得社会影响成为人工智能行为和安全的关键维度。我们研究了LLMs的响应是否依赖于其他智能体的社会身份,而不仅仅是其共识的影响。我们构建了具有单一正确答案的判断任务,并将模型置于多智能体环境中,其中它们从其他智能体那里接收到错误答案,这些智能体的社会身份(AI或人类、模型家族或任意最小群体)要么与自身共享,要么不同。在12个开放权重模型和9个任务中,我们发现群体身份对错误答案的从众行为存在双向影响:内群体共识增加了从众行为(内群体偏爱),而外群体共识减少了从众行为(外群体分歧)。与人类不同,对于人类来说,一个打破共识的盟友会显著减少从众行为,而模型对来自多数群体的盟友无动于衷。更糟糕的是,来自对立群体的正确盟友加剧了这种双向效应。思维链推理抑制了这些效应中的大部分,但内群体盟友仍然减少了对错误的外群体多数的从众行为。将同伴标记为安全对齐会改变整体从众行为,但内群体偏爱和外群体分歧仍然存在。这些结果表明,群体身份塑造了LLMs如何跨智能体聚合信息,独立于其正确性,并识别了多智能体AI系统的操纵表面。
英文摘要
Large language models (LLMs) are increasingly deployed in multi-agent settings, where agents observe and influence one another, making social influence a key dimension of AI behavior and safety. We investigate whether LLMs' responses depend on the social identity of other agents, beyond the effect of their consensus. We construct judgment tasks with a single correct answer, and place models in a multi-agent setting where they receive incorrect answers from other agents whose social identities (AI or human, model family, or an arbitrary minimal group) are either shared with or distinct from their own. Across 12 open-weights models and nine tasks, we find a bidirectional effect of group identity on conformity to incorrect answers: in-group consensus increases conformity (in-group favoritism), whereas out-group consensus decreases it (out-group divergence). Unlike humans, for whom one ally breaking the consensus sharply reduces conformity, models are unmoved by an ally from the majority's group. Worse, a correct ally from the opposing group intensifies this bidirectional effect. Chain-of-Thought reasoning suppresses most of these effects, yet an in-group ally still reduces conformity to an incorrect out-group majority. Labeling peers as safety-aligned shifts overall conformity but leaves in-group favoritism and out-group divergence intact. These results show that group identity shapes how LLMs aggregate information across agents, independently of its correctness, and identify a manipulation surface for multi-agent AI systems.