arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

以深度学习方式构建LLM智能体系统:从模块化设计到架构搜索

Building LLM Agent Systems the Deep Learning Way: From Modular Design to Architecture Search

Tao Feng, Pengrui Han, Zhongjie Dai, Jiaxuan You

arXiv 2610.04961首次发表:更新:

发表机构

University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

受深度学习启发,提出模块化构建LLM智能体系统,类比神经网络设计前向推理与反馈机制,并用架构搜索优化配置,实验显示性能显著提升。

AI 中文摘要

大型语言模型(LLMs)已经彻底改变了人工智能研究,并催生了令人兴奋的智能体系统。为了构建复杂的LLM智能体系统,大多数现有研究依赖于其他领域的见解或启发式方法,手动构建智能体系统。然而,这种方法通常需要大量的人工工程,并且无法完全优化感兴趣的下游任务。受深度学习巨大成功的启发,我们提出以模块化的方式构建LLM智能体系统,类似于构建深度神经网络。我们的关键见解是将LLM构建模块(如检索、记忆和提示策略)与成功的深度学习模块(如MLP、注意力和循环模块)进行类比。我们进一步为LLM设计了前向推理和反馈机制,其中LLM中的提示被视为深度模型中的权重,而来自反馈的提示优化类似于反向传播算法。我们还利用搜索算法来搜索LLM智能体系统的最佳配置,类似于深度学习研究中的神经架构搜索(NAS)。综合实验结果表明,所提出的用于LLM智能体系统的深度学习配方非常有效,特别是:(1)将LLM模块组织成深度学习风格的架构可带来显著的性能提升;(2)自动提示优化,等同于反向传播,在整合感兴趣任务的反馈方面效率高,并实现了至少5%的性能提升;(3)等同于NAS的算法在进一步优化LLM智能体系统架构方面效果良好,与随机设计的架构相比,性能提升了11%。总体而言,我们的研究展示了将深度学习的成功转移到构建LLM智能体系统的激动人心的机会。

英文摘要

Large Language Models (LLMs) have revolutionized AI research and enabled exciting agent systems. To build a complex LLM agent system, most existing research relies on insights from other domains or heuristics to manually build the agent system. However, this approach often requires heavy hand-engineering and fails to fully optimize for the downstream task of interest. Inspired by the tremendous success of deep learning, we propose to construct LLM agent systems in a modular manner, similar to building a deep neural network. Our key insight is to make analogies between LLM building blocks, such as retrievals, memories, and prompting strategies, and the successful deep learning modules, such as MLPs, attention, and recurrent modules. We further design forward inference and feedback mechanisms for LLMs, where prompts in LLMs are considered as the weights in deep models, and the prompt optimization from feedback is analogous to the back-propagation algorithm. We additionally leverage a search algorithm to search for the best configuration of LLM agent systems, similar to the neural architecture search (NAS) in deep learning research. Comprehensive experimental results demonstrate that the proposed deep learning recipe for LLM agent systems is highly effective, in particular: (1) Organizing LLM modules into deep-learning-style architectures yields noticeable performance gain; (2) Automatic prompt optimization, equivalent to backpropagation, is efficient in incorporating feedback from the task of interest and achieves at least 5% performance improvement; (3) NAS equivalent algorithm works well for further optimizing the LLM agent system architecture with 11% performance gain compared with randomly designed architectures. Overall, our research demonstrates the exciting opportunity of transferring the success of deep learning to building LLM agent systems.

Comments20 pages, 16 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑