发表机构
Peking University Shenzhen Graduate School(北京大学深圳研究生院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AgentFactory是同时优化智能体系统基础模型与工作流结构的框架,经8个跨域基准测试,其性能优于手动及现有自动化方法,平均提升9.1%。
AI 中文摘要
大型语言模型(LLMs)作为智能体系统中的强大组件,展现出了卓越的能力,能够实现复杂的推理和任务执行。然而,当前手动设计和优化智能体系统的方法严重依赖人工投入,限制了其适应性和可扩展性。近期的研究探索了工作流设计的自动化优化,但这些方法往往忽视了模型能力的关键作用,且仅关注单一性能指标,无法解决实际部署中的约束问题。在本文中,我们提出了AgentFactory,这一框架在智能体系统中同时优化基础模型和工作流结构,同时考虑性能、成本和效率等多个目标。AgentFactory利用先进的LLMs作为优化器,在庞大的可能配置搜索空间中进行探索,采用三阶段优化流程自动发现微调模型与优化工作流的有效组合。通过迭代优化过程,我们的框架系统地探索和评估不同的智能体系统设计,以适应特定任务需求,同时保持运行效率。我们在涵盖通用推理、编码、数学、医学和金融五个领域的八个基准上对AgentFactory进行评估。实验表明,AgentFactory在所有基准上的表现始终优于手动设计的方法和现有的自动化方法,平均提升达9.1%,在特定领域任务中提升尤为显著(MedQA提升19.6%,FinEval提升18.7%)。这些结果表明,AgentFactory是一种通过自动化优化开发更强大、高效智能体系统的有前景方法。
英文摘要
Large Language Models (LLMs) have demonstrated remarkable capabilities as powerful components in agentic systems, enabling sophisticated reasoning and complex task execution. However, current approaches to manually designing and optimizing agentic systems heavily rely on manual effort, limiting their adaptability and scalability. Recent work has explored the automated optimization of workflow designs. However, these approaches often overlook the crucial role of model capabilities and focus on single performance metrics, failing to address real-world deployment constraints. In this paper, we present AgentFactory, a framework that jointly optimizes both foundation models and workflow structures in agentic systems while considering multiple objectives including performance, cost, and efficiency. AgentFactory leverages advanced LLMs as optimizers to navigate the vast search space of possible configurations, employing a three-stage optimization pipeline to automatically discover effective combinations of fine-tuned models and optimized workflows. Through an iterative optimization process, our framework systematically explores and evaluates different agentic system designs, adapting to task-specific requirements while maintaining operational efficiency. We evaluate AgentFactory across eight benchmarks spanning five domains, including general reasoning, coding, mathematics, medicine, and finance. Our experiments demonstrate that AgentFactory consistently outperforms both manually designed methods and existing automated approaches, achieving an average improvement of 9.1% across all benchmarks, with particularly significant gains in domain-specific tasks (19.6% on MedQA and 18.7% on FinEval). These results establish AgentFactory as a promising approach for developing more capable and efficient agentic systems through automated optimization.