arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02758cs.MA

人人遵从,无人认同:LLM智能体群体中的多元无知

Everyone Conforms, No One Believes: Pluralistic Ignorance in LLM Agent Populations

Yashwanth YS

首次发表
浏览论文内容

中文总结 AI 辅助

该研究证实LLM智能体群体中会出现多元无知,构建了多场景基准评估8种模型,发现遵从率高且级联成功率低,提示模型选择会影响模拟结果,LLM模拟或高估社会规范稳定性。

中文摘要 AI 辅助

基于大语言模型(LLM)的多智能体系统越来越多地被用于模拟社会动态,从观点形成到集体决策,这些模拟能够复现某些社会现象,但尚不清楚它们是否捕捉到了多元无知——即多数人私下拒绝某一规范却公开遵从,且各自认为自己是唯一持异议者的状态。这一现象驱动着规范的持续存在、社会运动以及政治革命。我们证明,多元无知会在LLM智能体群体中稳健出现。我们基于人类多元无知文献构建了涵盖10个领域、5个权威层级的100个场景的基准,评估了来自6个机构的8种模型。尽管智能体私下反对该规范,其公开遵从率为64%至94%。遵从具有领域敏感性(职场和社会关系场景产生近乎普遍的遵从)且高度依赖模型,但与能力不相关。我们测试了单一“规范推动者”能否通过公开持异议打破虚假共识。在8种模型中,7种的级联成功率低于26%,其中一种模型在所有场景中均无级联;GPT-4o是显著异常值,达48%,揭示了不同模型家族间的定性动态差异。对所有8种模型的提示组件消融实验证实,遵从是涌现性的而非指令驱动的:移除虚假共识框架和适应目标会降低遵从但不会消除它(最小条件下为52%至92%)。我们的发现表明,模型选择是一种未被认可的自由度,从根本上塑造着模拟结果;更广泛而言,级联的近乎缺失表明,LLM模拟可能系统性高估社会规范的稳定性,忽略了驱动人类社会真实规范变革的脆弱临界点动态。

英文摘要

LLM-based multi-agent systems are increasingly used to simulate social dynamics, from opinion formation to collective decision-making. These simulations can reproduce certain social phenomena, but it is unknown whether they capture pluralistic ignorance, a state where a majority privately rejects a norm yet publicly conforms, each believing they are alone in dissenting. This phenomenon drives norm persistence, social movements, and political revolutions. We show that pluralistic ignorance emerges robustly in LLM agent populations. We construct a benchmark of 100 scenarios across 10 domains and 5 authority levels, grounded in the human pluralistic ignorance literature, and evaluate 8 models from 6 organizations. Agents publicly conform at rates of 64 to 94% despite privately opposing the norm. Conformity is domain-sensitive (workplace and social relationship scenarios produce near-universal compliance) and highly model-dependent, though uncorrelated with capability. We test whether a single "norm entrepreneur" can break the false consensus by publicly dissenting. For 7 of 8 models, cascades succeed less than 26% of the time, with one model showing zero cascades across all scenarios. GPT-4o is a notable outlier at 48%, revealing qualitatively distinct dynamics across model families. A prompt component ablation across all 8 models establishes that conformity is emergent rather than instruction-driven: removing both the false-consensus framing and fit-in goal reduces conformity but does not eliminate it (52 to 92% in the minimal condition). Our findings identify model selection as an unacknowledged degree of freedom that fundamentally shapes simulation outcomes. More broadly, the near-absence of cascades suggests LLM simulations may systematically overestimate the stability of social norms, missing the fragile tipping-point dynamics that drive real-world norm change in human societies.

↑