arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大规模交互语言模型群体中的共识与事实动态

Consensus and Factual Dynamics in Large Populations of Interacting Language Models

Emanuele Ricco, Elia Onofri, Vincenzo Sammartino, Roberto Di Pietro

arXiv 2609.39211首次发表:更新:

AI 中文总结

本研究提出物理启发的RHEON框架,将大规模LLM群体建模为自旋系统,通过扫描432种配置揭示共识涌现规律,发现交互拓扑影响收敛速度与幻觉最小化温度,且语义一致性与事实正确性正相关但不足以证明正确。

AI 中文摘要

大型语言模型(LLM)智能体正日益被部署为相互作用的实体群体,其中共识——对共享答案的一致意见——作为一种集体性的、非人为设计的涌现行为而出现。先前关于LLM共识的研究表明,智能体可以交叉验证彼此的答案并趋向于更符合事实的回应,将一致性视为正确性的代理指标。然而,这些研究通常固定单一交互结构,未阐明共识如何依赖于智能体之间的交互方式。我们通过引入RHEON框架来解决这一空白,该框架受物理学启发,将从一个冻结模型中抽取的群体重新诠释为一个在有效维度递增的交互几何阶梯上演化的$O(n)$自旋系统——从一维环到全耦合平均场图——以采样温度$T$作为热无序的可调来源,并通过类似Glauber的异步动力学进行演化。将RHEON扫描跨越432种提示、群体规模、通信拓扑和采样温度的配置组合,生成了Eraclitus-4.7M,一个包含470万条带标注的演化语料库。我们发现智能体在最初的几次更新扫描中达到最强的共识增益,且增加每个智能体的邻居数量平均而言会加速收敛。我们进一步表明,一个配置最终达成事实正确还是幻觉性共识,无法仅从其初始状态预测,并且最小化幻觉的温度取决于智能体的耦合方式,因此常见的近乎贪婪的默认设置并非自动最安全的选择。最后,语义一致性与事实收敛呈正相关,交互强化了这一关联,但共识程度永远不足以使一致性本身证明正确性。

英文摘要

Large Language Model (LLM) agents are increasingly deployed as populations of interacting entities, in which consensus --agreement on a shared answer-- emerges as a collective, unengineered behaviour. Prior work on LLM consensus shows that agents can cross-verify their answers and converge towards more factual responses, treating agreement as a proxy for correctness. However, these studies usually fix a single interaction structure, leaving open how consensus depends on how agents interact. We address this gap by introducing RHEON, a physics-inspired framework that recasts a population drawn from a single frozen model as an evolving $O(n)$ spin system on a ladder of interaction geometries of increasing effective dimension --from a 1D ring to a full-coupling mean-field graph-- with the sampling temperature $T$ as the tunable source of thermal disorder, evolved through a Glauber-like asynchronous dynamics. Sweeping RHEON across $432$ configurations of prompt, population size, communication topology, and sampling temperature yields Eraclitus-4.7M, a tagged evolutionary corpus of $4.7$ million responses. We find that agents reach their strongest consensus gain within the first few update sweeps and that increasing the number of neighbours per agent accelerates convergence on average. We further show that whether a configuration settles on factually correct or hallucinated consensus is not predictable from its initial state alone, and that the hallucination-minimising temperature depends on how the agents are coupled, so the common near-greedy default is not automatically the safest. Finally, semantic agreement correlates positively with factual convergence, and interaction strengthens the association, yet never enough for unanimity to certify correctness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑