发表机构
National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DeOrch通过将规划与工作者选择解耦,实现条件信用分配与在线适配,在减少调用次数的同时提升多智能体编排性能与迁移能力。
AI 中文摘要
学习式编排能够自动构建有效的语言模型多智能体系统,但现有方法将规划与固定的工作池耦合,并从相同的最终结果中训练分解与协作,这限制了迁移能力并模糊了信用分配。我们提出DeOrch,将不依赖工作者的规划与具体工作者选择相分离。其两阶段规划器首先在不含工作者信息的情况下分解任务,然后利用来自工作池的紧凑且不依赖工作者身份的匹配性反馈选择协作操作,从而实现对分解与协作决策的条件式信用分配。一个轻量级匹配器通过固定探针集上的行为估计工作者适用性,并借助上下文赌博机在线适应,使得新工作者无需重新训练规划器或匹配器即可被纳入。在多种分布内与分布外任务中,DeOrch在工作者调用次数少于竞争性学习编排器的情况下优于先前的自动多智能体系统编排方法,在迁移到完全未见的工作池且无需重新训练时仍保持有效,并显示出两个组件带来的持续增益。
英文摘要
Learned orchestration can automatically construct effective language-model multi-agent systems, but existing approaches couple planning to fixed worker pools and train decomposition and collaboration from the same terminal outcome, limiting transfer and obscuring credit assignment. We introduce DeOrch, which separates worker-agnostic planning from concrete worker selection. Its two-stage planner first decomposes the task without worker information, then chooses collaboration operations using compact, worker-identity-free matchability feedback from the pool, enabling conditional credit assignment to decomposition and collaboration decisions. A lightweight matcher estimates worker suitability from behavior on a fixed probe set and adapts online with a contextual bandit, allowing new workers to be incorporated without retraining the planner or matcher. Across diverse in- and out-of-distribution tasks, DeOrch outperforms prior automatic MAS orchestration methods with fewer worker calls than competing learned orchestrators, remains effective when transferred to an entirely unseen worker pool without retraining, and shows consistent gains from both components.