AI 中文总结
该研究针对多智能体LLM辩论中集体偏向性共识的涌现问题,提出物理启发的社会动力学分析框架,经实验验证其相变规律,发现智能体异质性可抑制该现象,且见解可推广至投资等现实决策任务。
AI 中文摘要
多智能体LLM辩论在决策任务和问题解决基准上表现出色,但其安全性与公平性风险仍未被充分理解。值得注意的是,交互作用会放大单个LLM的偏向性,引发了对实际部署的担忧。我们在多智能体LLM辩论中识别出集体(常具偏向性)规范的涌现,并表明噪声(如LLM采样温度)是关键驱动因素。为解释这一现象,我们提出了一个借鉴社会动力学的物理启发理论模型的分析框架。我们预测,当符合度超过由LLM初始偏向性和辩论噪声给定的临界阈值时,会发生向集体偏向性的相变。我们通过受控实验测试了这些理论预测,观察到与潜在相变一致的有限尺寸交叉现象。我们进一步发现,智能体异质性通过平滑( rounding)这一转变抑制了该涌现。最后,我们表明这些见解可推广到现实决策任务,包括投资决策和LLM作为评判者的评估。
英文摘要
Multi-agent LLM debates achieve strong performance on decision-making tasks as well as problem-solving benchmarks, yet their safety and fairness risks remain poorly understood. Notably, interaction can amplify the biases of single LLMs, raising concerns for real-world deployment. We identify the emergence of collective (often biased) norms in multi-agent LLM debates and show that noise (e.g., LLM sampling temperature) is a key driver. To explain this, we propose an analytical framework drawing on physics-inspired theoretical models of social dynamics. We predict a phase transition to collective bias when conformity surpasses a critical threshold given the LLMs' initial bias and debate noise. We test the theoretical predictions through controlled experiments and observe a finite-size crossover consistent with an underlying phase transition. We further find that agent heterogeneity suppresses emergence by smoothing (rounding) this transition. Finally, we show that these insights generalize to realistic decision-making tasks, including investment decisions and LLM-as-a-judge evaluation.
CommentsAccepted at the 43rd International Conference on Machine Learning (ICML 2026). 23 pages, 12 figures