发表机构
William & Mary; NEC Corporation of America; University of Illinois Urbana-Champaign(威廉与玛丽学院; 美国NEC公司; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多智能体工作流优化中仅训练生成器而忽略其他智能体的问题,提出FloWright框架,通过分层结构感知奖励实现角色自我进化与共同进化,并辅以DataWright数据强化,在多项任务上显著提升性能。
AI 中文摘要
处理复杂的现实世界任务可能超出单个大型语言模型(LLM)的能力,这促使人们使用多智能体工作流,通过协调专门的智能体共同处理这些任务。近期方法训练LLM根据执行结果构建更好的工作流,但它们仅优化工作流生成器,而构建或执行每个工作流的其他智能体保持固定,尽管每个结果都依赖于所有智能体。然而,将训练扩展到生成器之外具有挑战性:智能体是耦合的,且工作流的结果是一个单一的稀疏分数,无法指出哪个智能体导致了失败。我们提出FloWright,利用工作流作为优化工作流的框架。通过引入分层、结构感知的奖励范式,FloWright使一个角色能够自我进化,两个或更多角色能够共同进化,无需额外的模型、标签或执行。考虑到工作流通常在单个智能体已能处理的数据上进行训练和评估的局限性,我们进一步提出DataWright,一种自适应数据强化方法,将现有数据集转换为难度增加的工作流级任务。在文档、幻灯片、图表、代码、数学和金融任务中,使用FloWright训练的小型开放模型性能提升高达+7.41%,其中共同进化(+5.03%)比单独优化一个角色(+2.83%)获得更多收益。我们的项目页面:此https URL。
英文摘要
Tackling complex real-world tasks can exceed the capabilities of a single large language model (LLM), motivating the use of multi-agent workflows that coordinate specialized agents to work together on these tasks. Recent methods train LLMs to construct better workflows from execution outcomes, but they optimize only the workflow generator, while the other agents that build or execute each workflow remain fixed even though every outcome depends on all of them. However, extending training beyond the generator is challenging: the agents are coupled, and a workflow's outcome is a single sparse score that cannot tell which agent causes a failure. We propose FloWright, which leverages the workflow as a harness to optimize workflows. By introducing a hierarchical, structure-aware reward paradigm, FloWright enables one role to self-evolve and two or more roles to co-evolve, with no additional models, labels, or executions. Considering the limitation that workflows are commonly trained and evaluated on data that a single agent can already handle, we further propose DataWright, an adaptive data hardening approach that converts existing datasets into workflow-level tasks with increased difficulty. Across document, slide, chart, code, math, and finance tasks, small open models trained with FloWright achieve improved performance by up to $+7.41\%$, with co-evolving ($+5.03\%$) more roles gaining more than optimizing one of them alone ($+2.83\%$). Our project page: https://xhguo7.github.io/FloWright/.