融合即新变异:基于Bandit的工作流图进化
Fusion is the New Mutation: Bandit-Guided Evolution on Workflow Graphs
- School of Data Science, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院)
- College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
- Huawei Technologies Co., Ltd.(华为技术有限公司)
- Artificial Intelligence Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能学域)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出DAGO框架,利用上下文Bandit引导多父本融合优化智能体工作流,在有限评估预算下高效探索,在六个基准上优于基线,并提升性能与降低开销。
AI中文摘要:
自动化智能体工作流优化依赖于昂贵的评估,因此有效分配有限的评估预算至关重要。多父本融合可以重用先前发现的工作流中的设计,但识别有前景的父本组合需要从有限的融合反馈中学习。我们提出DAGO(有向无环图优化),一种上下文Bandit引导的框架,在有限的评估预算下学习融合哪些父本工作流。DAGO将每个候选父本组合视为一个臂,由其所含工作流的代码和提示的预训练嵌入表示。对角LinUCB策略学习跨臂的共享奖励模型,并在具有高预测后代质量的臂的利用与基于不确定性的探索之间取得平衡。选择臂后,LLM通过摘要引导的融合生成子工作流,子工作流的验证分数作为更新Bandit的奖励。共享的有向无环图维护已发现的工作流及其多父本谱系,为后续臂提议提供不断扩大的父本池。在涵盖数学推理、代码生成和问答的六个基准上,DAGO在评估的基线中取得了最高的宏观平均分数。在匹配的验证-评估预算下,它相对于AFlow从80.3提高到81.7,同时将总搜索开销减少11.2%。消融研究表明,LinUCB引导的臂选择优于随机选择及其无探索变体,支持反馈驱动选择和探索-利用平衡的价值。
英文摘要:
Automated agentic workflow optimization relies on costly evaluations, making it essential to allocate a limited evaluation budget effectively. Multi-parent fusion can reuse designs from previously discovered workflows, but identifying promising parent combinations requires learning from limited fusion feedback. We introduce DAGO (Directed Acyclic Graph Optimization), a contextual-bandit-guided framework that learns which parent workflows to fuse under a limited evaluation budget. DAGO formulates each candidate parent combination as an arm, represented by pretrained embeddings of its constituent workflows' code and prompts. A diagonal LinUCB policy learns a shared reward model across arms and balances exploitation of arms with high predicted offspring quality against uncertainty-driven exploration. After an arm is selected, an LLM generates a child workflow through summary-guided fusion, and the child's validation score serves as the reward for updating the bandit. A shared directed acyclic graph maintains discovered workflows and their multi-parent lineage, providing an expanding pool of parents for subsequent arm proposals. Across six benchmarks covering mathematical reasoning, code generation, and question answering, DAGO achieves the highest macro-average score among the evaluated baselines. Under matched validation-evaluation budgets, it improves over AFlow from 80.3 to 81.7 while reducing aggregate search expenditure by 11.2%. Ablation studies show that LinUCB-guided arm selection outperforms both random selection and its exploration-free variant, supporting the value of feedback-driven selection and exploration-exploitation balance.