arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35885cs.MA

多智能体LLM中的集体状态:推理努力与通信拓扑的影响

Collective Regimes in Multi-Agent LLMs under Reasoning Effort and Communication Topology

  • University of Pennsylvania(宾夕法尼亚大学)
  • Carnegie Mellon University(卡内基梅隆大学)
  • Shibaura Institute of Technology(芝浦工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Machiko Hirota, Akshara Nadayanur Sathis Kanna, Ujwal Kumar, Phan Xuan Tan

AI总结:

本研究通过N=50个无状态LLM智能体实验,识别出同步、扭曲和类嵌合体三种集体状态,发现推理努力促进局部有序,而通信连通性驱动全局同步,且总体一致性不足以表征集体行为。

AI中文摘要:

多智能体LLM系统越来越多地被用于审议和评估,通常假设更多的同伴互动会导致更可靠的共识。现有工作主要通过最终准确性或总体一致性来评估这些系统。然而,这些措施并不能揭示一致性在专家组中是如何组织的。在本文中,我们研究了N=50个无状态LLM智能体,它们从局部可见的同伴更新预测,并使用全局和局部的一致性测量来表征其行为。我们识别出三种集体状态:同步、扭曲(局部有序但全局不一致)和类嵌合体,其中一致和不一致的子群体共存。增加gpt-5-mini中的推理努力使专家组从多变、通常碎片化的结果转向局部有序的扭曲状态,一个小型后续研究表明,这种状态也可以从置换的初始条件形成,而增加通信连通性则推动它们走向全局同步。随着重连图中代数连通性的增加,碎片化崩溃得更快。拓扑效应也出现在非循环评判任务上,并跨越三家提供商的模型。最后,低空间异质性并不能保证全局共识:在ΔZ低于0.03的试验中,40%在最后20轮中保持扭曲配置。这些结果表明,推理努力和通信拓扑控制多智能体协调的不同方面,仅凭总体一致性不足以表征集体LLM行为。

英文摘要:

Multi-agent LLM systems are increasingly used for deliberation and evaluation, often under the assumption that greater peer interaction leads to more reliable consensus. Existing work largely evaluates these systems through final accuracy or aggregate agreement. However, such measures do not reveal how agreement is organized in the panel. In this paper, we study N = 50 stateless LLM agents that update their predictions from locally visible peers, and characterize their behavior using both global and local measurements of agreement. We identify three collective regimes: synchronised, twisted (locally ordered but globally incoherent) and chimera-like, where coherent and incoherent subpopulations coexist. Increasing reasoning effort in gpt-5-mini shifts panels from variable, often fragmented outcomes toward locally ordered twisted states, and a small follow-up shows such states can also form from permuted initial conditions, whereas increasing communication connectivity drives them toward global synchronisation. Fragmentation collapses faster as algebraic connectivity increases across rewired graphs. The topology effect also appears on a non-circular judging task and across models from three providers. Finally, low spatial heterogeneity does not guarantee global consensus: 40% of trials with Delta Z below 0.03 retain a twisted configuration through the final 20 turns. These results show that reasoning effort and communication topology control different aspects of multi-agent coordination, and that aggregate agreement alone is insufficient to characterize collective LLM behavior.

↑