arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为故障付费,而非流程:无标签的流内多智能体工作流优化

Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization

Xuehang Guo, Haoyu Wang, Shengyu Chen, Zach Chen, Wei Cheng, Qingyun Wang, Haifeng Chen

arXiv 2610.01017首次发表:更新:

发表机构

William & Mary; NEC Corporation of America(威廉与玛丽学院; 美国NEC公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出InFlowOp,一种无标签的流内多智能体工作流优化方法,通过统一成本定价决策,在执行前双向确定任务分解与智能体分配,执行中低成本纠错,在Braid基准上超越单智能体基线最高11.97%。

AI 中文摘要

大型语言模型(LLMs)越来越多地构建多智能体工作流,将复杂任务分解并从智能体池中分配专家智能体。然而,构建这样一个工作流仍然具有挑战性:任务划分的精细程度、每个子任务信任哪个智能体、以及何时创建新的专家,都是工作流构建者需要预先确定的关键决策。因此,每个子任务是否成功在工作流运行之前是未知的。然而,改进工作流成本高昂。定位故障通常需要参考答案、分级结果或训练好的评估器,并且修复措施通过重新执行、重新搜索或重新训练应用于整个工作流。我们提出InFlowOp,它用单一无标签成本为每个决策定价,该成本衡量智能体的能力与子任务需求的匹配程度,以及该智能体运行所需的成本。在执行前,InFlowOp根据成本而非固定模板双向确定任务分解的粒度和智能体分配。在执行期间,InFlowOp通过相同的成本以最便宜的移动纠正故障,该成本在工作流构建和运行时均适用。面对工作流级评估的挑战,我们引入Braid,一个其任务需要超越单智能体能力的多智能体协调的基准。在各种领域和骨干网络上,InFlowOp优于单智能体基线,最高提升+11.97%,在流内优化下达到+9.64%。我们的项目页面:此https URL。

英文摘要

Large language models (LLMs) increasingly construct multi-agent workflows that decompose a complex task and assign specialist agents from a pool. However, building such a workflow well remains challenging: how finely to divide the task, which agent to trust with each subtask, and when to create a new specialist are all critical decisions a workflow constructor needs to settle up front. Thus, whether each subtask succeeds remains unknown until the workflow runs. Yet, improving a workflow is costly. Locating a fault usually requires a reference answer, a graded outcome, or a trained assessor, and the fix is applied to the whole workflow through re-execution, re-search, or retraining. We propose InFlowOp, which prices every decision in one label-free cost that weighs how well an agent's competence meets what a subtask demands against how much that agent takes to run. Before execution, InFlowOp bidirectionally determines the granularity of task decomposition and agent assignment following from the cost rather than from a fixed template. During execution, InFlowOp corrects a fault with the cheapest move via the same cost that serves the workflow both as it is built and as it runs. Facing the workflow-level evaluation challenge, we introduce Braid, a benchmark whose tasks require multi-agent coordination beyond single-agent capability. Across various domains and backbones, InFlowOp outperforms single agent baselines by up to $+11.97\%$, achieving $+9.64\%$ with in-flow optimization. Our project page: https://xhguo7.github.io/InFlowOp/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑