CollabFlow:智能体协作的递归自我改进
CollabFlow: Recursive Self-Improvement of Agent Collaboration
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Fudan University(复旦大学)
- University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多智能体协作中协作预定义、错误传播和团队集中问题,提出CollabFlow,通过可训练的导演构建团队并基于证据条件通信和协作轨迹平衡目标实现递归自我改进,在十二个数据集上优于基线。
AI中文摘要:
递归自我改进(RSI)使系统能够从自身的结果中改进;在基于LLM的多智能体系统中,智能体在任务内相互改进,而结果则改善它们跨任务的协作方式。然而,现有的多智能体协作使这个循环保持开放:协作在操作员层面被预定义,仅拓扑学习保留逐字交换从而传播错误,并且对系统自身结果的奖励最大化集中于少数团队。为应对这些挑战,我们提出CollabFlow,一个学习型智能体协作的RSI系统:一个可训练的Collab-Director构建完整智能体的团队,一个冻结的执行器运行它们,每轮的结果重新训练该导演。在每一轮内,协作图的边携带证据条件通信的协议:接收方仅在发送方的证据强度超过一定阈值时才采纳不同的答案,因此导演学习谁通信以及如何通信。跨轮次,我们进一步提出协作轨迹平衡(CTB),一种基于流的目标,它在构建顺序上对每个团队只记一次功劳,并针对团队上的奖励比例分布,从而让多个优秀团队保持活跃。我们还限制了这种自生成目标在轮次之间的移动距离,随着记录的积累而缩小。在十二个数据集上,CollabFlow优于所有基线并跨轮次持续改进。代码可在该https URL获取。
英文摘要:
Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing multi-agent collaboration leaves this loop open: collaboration is pre-defined at the operator level, topology-only learning keeps verbatim exchange that propagates errors, and reward maximization on a system's own outcomes concentrates on a few teams. To address these challenges, we propose CollabFlow, an RSI system of Learned Agent Collaboration: a trainable Collab-Director constructs teams of complete Agents, a frozen executor runs them, and each round's outcomes retrain the director. Within each round, the edges of a collaboration graph carry protocols of Evidence-Conditioned Communication: a receiver adopts a differing answer only when the sender's evidence is stronger by a margin, so the director learns who communicates and how. Across rounds, we further propose Collaborative Trajectory Balance (CTB), a flow-based objective that credits each team once across its construction orders and targets a reward-proportional distribution over teams, so several good teams stay in play. We also bound how far this self-generated target moves between rounds, which shrinks as records accumulate. On twelve datasets, CollabFlow outperforms all baselines and keeps improving across rounds. Code is available at https://anonymous.4open.science/r/CollabFlow-631E.