发表机构
Shanghai Academy of AI for Science (SAIS); AI Institute, Fudan University; Yunnan University(上海人工智能科学研究院; 复旦大学人工智能研究院; 云南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Waggle,一种可共享的匿名局部法则,通过群体一致蒸馏学习,使LLM智能体群无需显式角色或全局拓扑即可自组织协调,保留96%以上神谕质量并可迁移。
AI 中文摘要
随着LLM智能体在复杂任务上的协作日益增多,如何组织它们的交互成为一个核心设计问题。现有的多智能体系统通常学习或适应显式的角色、层级、路由策略或通信拓扑。我们将学习目标转向一个可复用的局部法则,该法则可在可互换的智能体之间共享,并随着群体或交互条件的变化调整协调,而无需重新定义全局组织。我们提出了Waggle,一种基于有界局部视图的共享匿名策略,它联合选择任务动作、语义通信和局部承诺更新。相同法则的重复执行使得协调能够在没有显式角色或全局拓扑的情况下在线形成、持续和重组。为了在可互换的智能体和演化的协调中学习这一法则,我们开发了群体一致蒸馏(SCD),将匿名轨道一致性与基于rollout的下一局部协调场预测相结合,且无需增加推理时组件。在多种协调设置中,同一学习到的法则在群体和交互预算变化时仍然有效,保留了超过96%的特定基底神谕质量,并且无需重新训练即可迁移;SCD进一步改善了反证后的重组。这些结果共同表明,LLM智能体的组织可以通过重复执行学习到的局部法则而涌现和适应。
英文摘要
As LLM agents increasingly collaborate on complex tasks, how to organize their interactions becomes a central design question. Existing multi-agent systems typically learn or adapt explicit roles, hierarchies, routing policies, or communication topologies. We shift the learning target to a reusable local law that can be shared across interchangeable agents and adapt coordination as populations or interaction conditions change, without redefining a global organization. We introduce Waggle, a shared anonymous policy over bounded local views that jointly selects task actions, semantic communication, and local commitment updates. Repeated execution of the same law allows coordination to form, persist, and reorganize online without explicit roles or global topology. To learn this law across interchangeable agents and evolving coordination, we develop Swarm-Consistent Distillation (SCD), combining anonymous-orbit consistency with rollout-grounded prediction of the next local coordination field, with no added inference-time components. Across diverse coordination settings, the same learned law remains effective as populations and interaction budgets change, retains over 96% of substrate-specific oracle quality, and transfers without retraining; SCD further improves reorganization after counterevidence. Together, these results show that LLM-agent organization can emerge and adapt through repeated execution of a learned local law.