发表机构
William & Mary; NEC Corporation of America(威廉与玛丽学院; 美国NEC公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出InFlowOp,一种无标签的流内多智能体工作流优化方法,通过统一成本定价决策,在执行前双向确定任务分解与智能体分配,执行中低成本纠错,在Braid基准上超越单智能体基线最高11.97%。
AI 中文摘要
大型语言模型(LLMs)越来越多地构建多智能体工作流,将复杂任务分解并从智能体池中分配专家智能体。然而,构建这样一个工作流仍然具有挑战性:任务划分的精细程度、每个子任务信任哪个智能体、以及何时创建新的专家,都是工作流构建者需要预先确定的关键决策。因此,每个子任务是否成功在工作流运行之前是未知的。然而,改进工作流成本高昂。定位故障通常需要参考答案、分级结果或训练好的评估器,并且修复措施通过重新执行、重新搜索或重新训练应用于整个工作流。我们提出InFlowOp,它用单一无标签成本为每个决策定价,该成本衡量智能体的能力与子任务需求的匹配程度,以及该智能体运行所需的成本。在执行前,InFlowOp根据成本而非固定模板双向确定任务分解的粒度和智能体分配。在执行期间,InFlowOp通过相同的成本以最便宜的移动纠正故障,该成本在工作流构建和运行时均适用。面对工作流级评估的挑战,我们引入Braid,一个其任务需要超越单智能体能力的多智能体协调的基准。在各种领域和骨干网络上,InFlowOp优于单智能体基线,最高提升+11.97%,在流内优化下达到+9.64%。我们的项目页面:此https URL。
英文摘要
Large language models (LLMs) increasingly construct multi-agent workflows that decompose a complex task and assign specialist agents from a pool. However, building such a workflow well remains challenging: how finely to divide the task, which agent to trust with each subtask, and when to create a new specialist are all critical decisions a workflow constructor needs to settle up front. Thus, whether each subtask succeeds remains unknown until the workflow runs. Yet, improving a workflow is costly. Locating a fault usually requires a reference answer, a graded outcome, or a trained assessor, and the fix is applied to the whole workflow through re-execution, re-search, or retraining. We propose InFlowOp, which prices every decision in one label-free cost that weighs how well an agent's competence meets what a subtask demands against how much that agent takes to run. Before execution, InFlowOp bidirectionally determines the granularity of task decomposition and agent assignment following from the cost rather than from a fixed template. During execution, InFlowOp corrects a fault with the cheapest move via the same cost that serves the workflow both as it is built and as it runs. Facing the workflow-level evaluation challenge, we introduce Braid, a benchmark whose tasks require multi-agent coordination beyond single-agent capability. Across various domains and backbones, InFlowOp outperforms single agent baselines by up to $+11.97\%$, achieving $+9.64\%$ with in-flow optimization. Our project page: https://xhguo7.github.io/InFlowOp/.