arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33871cs.MAcs.CL

群体物理学,群体问题:LLM 社会中的安全性与涌现性

Population Physics, Population Problems: Safety and Emergence in LLM Societies

Adrian de Wynter

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出衡量LLM社会自组织的框架,应用于三个系统发现显著自组织现象,并揭示群体病态可由子集协调驱动,为多智能体系统提供轻量级诊断方法。

中文摘要 AI 辅助

大型语言模型(LLM)社会的集体行为并非其个体输出的简单总和。它会产生统计上显著不同、有时不可预测的现象,而我们用于研究单个智能体的工具可能无法扩展到这些现象。然而,由于近期涉及自主智能体系统的事件,理解这些系统至关重要。为此,我们引入了一个用于衡量LLM社会系统中自组织现象的框架,并将其应用于三个此类系统:一个谢林网格、一个社交网络(Moltbook)以及一个类似Twitter的虚假信息模拟系统('Rogue')。这三个系统均展现出统计上显著的自组织现象。此外,它们的弛豫动力学随智能体可获取的环境信息而变化,其中开放式系统(Moltbook、Rogue)表现出尖锐的、类似相变的动力学特征。进一步的结果表明,即使LLM经过安全调优或受到监控,群体层面的病态现象也可能出现,这主要由群体子集的协调活动驱动。我们还展示了在另外两种场景(一个公共地困境,GovSim,以及一个LLM作为法官的审议方案,ChatEval)下自组织现象并未出现。我们认为,测量此类特征提供了一种轻量级、与智能体无关的诊断层,用于检测部署的多智能体系统中的协调集体行为,而无需依赖自然语言或模型版本信息。

英文摘要

The collective behaviour of large language model (LLM) societies is not the sum of their individual outputs. It yields statistically distinct, sometimes-unpredictable phenomena, for which the tools we use to study single agents may not scale. Due to recent incidents involving autonomous agentic systems, however, understanding these systems is paramount. For that we introduce a framework for measuring self-organisation in LLM social systems and apply it to three such systems: a Schelling grid, a social network (Moltbook), and a Twitter-like misinformation simulation ('Rogue'). All three exhibit statistically significant self-organisation. Moreover, their relaxation dynamics vary with the environmental information available to the agents, with open-ended systems (Moltbook, Rogue) exhibiting sharp, phase-transition-like dynamics. Further results show that population-level pathologies can emerge even when the LLMs are safety-tuned or monitored, being primarily driven by the coordinated activity of a population subset. We also show when self-organisation does \textit{not} emerge under two additional scenarios (a commons dilemma, GovSim, and a LLM-as-a-judge deliberation scheme, ChatEval). We argue that measuring signatures of this kind offers a lightweight, agent-agnostic diagnostic layer for detecting coordinated collective behaviour in deployed multi-agent systems without relying on natural language or model versioning.

发表机构

  • Microsoft(微软)
  • The University of York(约克大学)

机构由 AI 辅助整理,请以论文原文为准。

↑