场景理论:打破多智能体LLM协调中的对称性陷阱
Theory of Scene: Breaking the Symmetry Trap in Multi-Agent LLM Coordination
AI总结:
针对多智能体LLM协调中的对称性陷阱,提出免训练的场景理论(ToS),通过角色门控和任务耦合实现同质智能体分工,在DivvyBench、GovSim和Overcooked上全面超越六个基线。
AI中文摘要:
基于大型语言模型(LLM)构建的多智能体系统在很大程度上是同质的,因为即使使用不同的LLM,其智能体行为也趋于一致。我们表明,当此类智能体在没有通信的情况下并发行动时,它们会在必须分工的目标上发生冲突,而在必须共同行动的目标上发生分歧,这种双重失败我们称之为对称性陷阱。心理理论(ToM)被广泛用于无通信协调,但无法摆脱这一陷阱,因为同质智能体会对彼此形成相同的预测,并以相同方式响应这些预测。我们提出场景理论(ToS),一种免训练的推理模式,其中每个智能体读取其公共角色——智能体之间唯一的差异——以及它们共同观察的任务上下文。同质智能体由此推导出单一的分工方案,每个智能体承担其角色所固定的部分,这将同质性从陷阱的成因转变为解药。ToS通过角色门控将角色与场景一起读取,角色门控决定所有权是重叠还是已经划分;通过任务耦合读取任务上下文,任务耦合推断团队必须对每个目标收敛、划分目标,还是依次分阶段行动。我们在DivvyBench上评估了该方法,这是一个我们引入的受控环境,其目标类型使一个回合在桌面、空域和家庭场景中成为竞争型、合作型或混合型;并在两个既有的智能体基准上评估,即GovSim和Overcooked。ToS在每个基准上都优于所有六个基线,且每个基线至少在一种设置中远远落后于ToS。与给定相同输入的心理理论相比,ToS将DivvyBench成功率从71.1%提高到99.6%,将GovSim总收益从207提高到400,并将Overcooked水平归一化吞吐量从1.41提高到1.67。
英文摘要:
Multi-agent systems built on large language models (LLMs) are largely homogeneous, as their agents behave alike even across distinct LLMs. We show that when such agents act concurrently without communication, they collide on targets they must split and diverge on targets they must take together, a double failure we term the symmetry trap. Theory of Mind (ToM), widely used for coordination without communication, cannot escape this trap, since homogeneous agents form the same prediction of one another and respond to it in the same way. We propose Theory of Scene (ToS), a training-free reasoning schema in which each agent reads its public role, the only difference between the agents, and the task context they all observe. Homogeneous agents thereby derive one division of labor, each taking the part its role fixes, which turns homogeneity from the cause of the trap into the cure. ToS reads the role together with the scene through role gating, which determines whether ownership overlaps or is already divided, and the task context through task coupling, which infers whether the team must converge on each target, divide it, or take its stages in turn. We evaluate on DivvyBench, a controlled environment we introduce, whose target types make an episode Competitive, Cooperative, or Mixed across Tabletop, Airspace, and Household scenarios, and on two established agentic benchmarks, GovSim and Overcooked. ToS outperforms all six baselines on every benchmark, and each baseline falls far behind it in at least one setting. Against ToM given the same inputs, ToS raises the DivvyBench success rate from 71.1% to 99.6%, the GovSim total gain from 207 to 400, and the Overcooked level-normalized throughput from 1.41 to 1.67.