arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MoFlow:多目标智能体工作流生成

MoFlow: Multi-Objective Agentic Workflow Generation

Yining Lu, Aurelie Lozano, Xi Yang, Naoki Abe, Yu Deng, Meng Jiang

arXiv 2609.38294首次发表:更新:

发表机构

University of Notre Dame; IBM(圣母大学; IBM)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MoFlow将智能体工作流生成建模为多目标马尔可夫决策过程,利用凸包蒙特卡洛树搜索覆盖帕累托前沿,实现任意偏好下无需重训的工作流生成,并在六个基准上取得最高平均超体积。

AI 中文摘要

我们研究智能体工作流的生成问题,该问题需要同时优化多个目标,如准确性、成本、延迟、鲁棒性和一致性。现有的工作流生成方法通常仅优化准确性或目标的加权和,因此每个训练好的生成器都固定于一种权衡,当偏好改变时,必须从头重新训练。为缓解这一问题,我们提出MoFlow,它生成针对不同偏好进行优化的工作流。具体而言,MoFlow将工作流生成形式化为多目标马尔可夫决策过程,并通过利用带有乐观集值备份的凸包蒙特卡洛树搜索来求解,其中每个节点存储一组可达的权衡而非单一的加权分数。因此,单次搜索即可近似覆盖帕累托前沿,MoFlow可通过查表为任何偏好返回工作流而无需重新训练。我们在涵盖数学、代码和问答的六个基准上,将MoFlow与六个强基线进行了评估。由于基线在设计上是单标量优化器,进行苹果对苹果的比较较为困难。我们转而采用一种有利于基线的评估设置,即针对每个测试偏好重新运行基线,而MoFlow从未见过这些偏好。即使在这种严格的设置下,MoFlow仍取得了最高的平均超体积。

英文摘要

We study the generation of agentic workflows that jointly optimize multiple objectives, such as accuracy, cost, latency, robustness, and consistency. Existing methods for workflow generation typically optimize accuracy alone or a weighted sum of objectives, so each trained generator commits to one fixed trade-off and must be retrained from scratch when preferences change. To alleviate this, we propose MoFlow, which generates workflows optimized across varied preferences. Specifically, MoFlow formulates workflow generation as a multi-objective Markov decision process and solves it by leveraging Convex-Hull Monte Carlo Tree Search with optimistic set-valued backups, where every node stores a set of reachable trade-offs rather than one weighted score. A single search thus approximately covers the Pareto front, from which MoFlow can return a workflow for any preference by lookup without retraining. We evaluate MoFlow against six strong baselines on six benchmarks spanning mathematics, code, and question answering. Since the baselines are single-scalar optimizers by design, an apples-to-apples comparison is difficult. We instead adopt an evaluation setup that favors the baselines, in that they are rerun for each testing preference, which MoFlow never sees. Even under this stringent setup, MoFlow achieves the highest average hypervolume.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑