发表机构
Southeast University; Rensselaer Polytechnic Institute; The Hong Kong University of Science and Technology(东南大学; 伦斯勒理工学院; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在让大型语言模型智能体通过工作流解决复杂任务。提出FlowEvo框架,通过工作流到技能编译、技能到工作流反馈、技能管理三个机制,使智能体无需更新模型参数积累完善能力,实验显示其在基准测试中精度-成本权衡优,各机制有贡献。
AI 中文摘要
大型语言模型智能体越来越多地通过构建结合推理、工具使用和代码执行的推理时工作流来解决复杂任务。虽然这些工作流能够灵活地解决问题,但执行过程中发现的有用过程往往是短暂的。我们提出了FlowEvo,一个无需训练的框架,它将成功的轨迹编译成可重复使用的技能记录。每个记录将一个可调用工件与辅助结构化指导配对,并在可行时进行接口、重放和安全检查。这些技能记录在推理时保存在技能库中。FlowEvo围绕三个耦合机制组织:工作流到技能的编译,从成功轨迹中提取可重复使用的可执行工件;技能到工作流的反馈,检索积累的技能以支持未来的问题解决;技能管理,监测下游效用并抑制导致负迁移的技能。通过这个工作流-技能-工作流反馈循环,FlowEvo使智能体能够随着时间的推移积累和完善任务解决能力,而无需更新模型参数。在跨越交互式环境(ALFWorld)和代码/数学生成(HumanEval、GSM8K)的基准测试中进行的实验表明,在我们的实现设置下,FlowEvo在评估的基线中实现了最佳精度-成本权衡。在ALFWorld上,FlowEvo实现了82.8%的成功率,比最强基线高出23.6个百分点,而其每集的平均令牌使用量不到最有效基线的一半。受控消融实验证实每个机制都对整体结果有贡献。代码可在该https URL上公开获取。
英文摘要
Large language model agents can adapt to complex tasks by constructing workflows at inference time, but procedures discovered in one episode are usually discarded after execution. Existing skill libraries provide reusable executable routines, but are typically assembled offline and do not grow from the agent's own workflows. We introduce FlowEvo, a training-free framework in which workflows and skills co-evolve at inference time. FlowEvo compiles successful workflows into callable skills, stores them in a persistent bank, and uses retrieved skills either through direct execution or as context for constructing new workflows. It also tracks each skill's downstream utility and suppresses skills that cause negative transfer. Using a shared GPT-4o-mini backbone, FlowEvo achieves the highest accuracy among 8 baselines on the full standard splits of ALFWorld, HumanEval, MBPP, GSM8K, and MATH-500. On ALFWorld, it reaches 85.6%, 26.4 points above the strongest baseline, while using roughly one third as many tokens. Across 10 base models spanning 7B to 671B parameters, FlowEvo outperforms ExpeL in 49 of 50 model-dataset comparisons. Code is available at https://github.com/DEFENSE-SEU/FlowEvo.
CommentsPublished as a conference paper at the Conference on Language Modeling (COLM) 2026. 25 pages, 3 figures, 16 tables. Code: https://github.com/DEFENSE-SEU/FlowEvo