arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FlowScout:从执行反馈到可靠的工具使用智能体工作流

FlowScout: From Execution Feedback to Reliable Tool-Using Agent Workflows

Shuo Hao, You Lu, Bihuan Chen, Xin Peng

arXiv 2608.10039首次发表:更新:

AI 中文总结

FlowScout是从历史任务记录生成工具集成型智能体工作流的执行引导框架,经多任务域评估,其工具调用正确性与执行质量均显著优于PM4Py等基线方法。

AI 中文摘要

智能体工作流已成为构建可靠的基于大语言模型(LLM)的自动化系统的重要抽象,它将LLM、工具和控制逻辑组织成显式的执行结构。然而,构建高质量的智能体工作流在很大程度上仍依赖人工,且需要大量领域专业知识。近期研究探索从历史任务解决记录中自动生成智能体工作流,但这些研究主要生成以LLM为中心的工作流,其中实际工具执行被LLM节点抽象和模拟,限制了生成工作流的可用性和稳定性。为解决这些局限,我们提出FlowScout,这是一个从历史任务解决记录中生成工具集成型智能体工作流的执行引导框架。具体而言,FlowScout将智能体工作流表示为由LLM节点、工具调用节点和依赖边组成的有向图。它首先从历史记录中挖掘通用工具协调骨架以构建初始工作流,随后在执行反馈的引导下通过蒙特卡洛树搜索优化工作流拓扑。我们在四个代表性任务领域对FlowScout进行评估,并与PM4Py、ReAct和AFlow三个基线方法对比。实验结果显示,FlowScout生成的智能体工作流较基线方法将工具调用正确性至少提升92.69%,执行质量至少提升17.66%,同时在重复运行中实现更低的性能波动。

英文摘要

Agentic workflows have become an important abstraction for building reliable LLM-based automation systems by organizing large language models (LLMs), tools, and control logic into explicit execution structures. However, constructing high-quality agentic workflows remains largely manual and requires substantial domain expertise. Recent studies have explored automatic agentic workflow generation from historical task-solving records, but they mainly produce LLM-centric workflows, where real tool executions are abstracted and simulated by LLM nodes, limiting the usability and stability of generated workflows. To address these limitations, we propose FlowScout, an execution-guided framework for generating tool-integrated agentic workflows from historical task-solving records. Specifically, FlowScout represents an agentic workflow as a directed graph composed of LLM nodes, tool-calling nodes, and dependency edges. It first mines a common tool coordination skeleton from historical records to construct an initial workflow, and then refines the workflow topology through Monte Carlo tree search guided by execution feedback. We evaluate FlowScout on four representative task domains and compare it with three baselines, i.e., PM4Py, ReAct and AFlow. Experimental results show that agentic workflows generated by FlowScout improve tool invocation correctness by at least 92.69% and execution quality by at least 17.66% over the baselines, while achieving lower performance variation across repeated runs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑