在软件工程中开发基于大语言模型(LLM)的多智能体系统:一项混合方法经验报告
Developing LLM-based Multi-Agent Systems in Software Engineering: A Mixed-Method Experience Report
浏览论文内容
中文总结 AI 辅助
本文通过定量和定性分析,梳理软件工程中基于LLM的MAS现有框架,发现其基本组件覆盖良好但缺高级功能,摘要任务ROUGE分数无显著差异,为相关人员选框架提供指导。
中文摘要 AI 辅助
由大语言模型(LLM)驱动的生成式人工智能(Gen AI)的普及,改变了软件开发流程,为代码生成、调试、测试和维护引入了新范式。早期应用聚焦于利用单一独立的LLM协助开发者完成孤立任务,而近期进展已转向多智能体系统(MAS),该系统协调多个基于LLM的智能体协同实现共同目标。尽管MAS前景广阔,但开发者使用时面临一系列挑战,包括需精心选择合适技术、设计恰当的协调规则、为涉及的智能体规划特定角色。本文全面概述了软件工程中用于实现MAS的现有工具与框架:首先,从开发者视角对最相关的开源MAS框架开展定量分析,评估其文档、特性与能力;其次,通过实现一个通用用例(.http URL文件的摘要生成),对部分选定框架进行定性评估。研究发现,选定框架对MAS的基本组件覆盖良好,但智能体遥测等高级功能仍有缺失;此外,针对摘要生成任务的经验评估显示,ROUGE分数无显著差异。最后,本文给出一系列经验教训与挑战,可帮助研究者和从业者根据自身需求选择合适的MAS框架。
英文摘要
The proliferation of Generative Artificial Intelligence (Gen AI) powered by large language models (LLMs) has transformed the software development process, introducing new paradigms for code generation, debugging, testing, and maintenance. While early applications focused on leveraging single, independent LLMs to assist developers with isolated tasks, recent advances have shifted toward multi-agent systems (MAS) that orchestrate multiple LLM-based agents working collaboratively toward common objectives. Despite their promising potential, using MAS encompasses a set of challenges for developers who have to carefully select the right technology, devise proper coordination rules, and design specific roles for the involved agents. In this paper, we provide a comprehensive overview of the existing tools and frameworks for implementing MAS in software engineering. First, we conducted a quantitative analysis of the most relevant open source MAS frameworks by evaluating their documentation, features, and capabilities from the developers' perspective. Second, we performed a qualitative evaluation of a subset of the selected frameworks by implementing a common use case: the summarization of README.MD files. The findings show that the selected frameworks provide a good coverage of fundamental components of MAS, though advanced features such as telemetry of agents are still missing. In addition, the empirical evaluation shows that there is no significant difference in terms of ROUGE scores considering the summarization task. Finally, we provide a set of lessons learned and challenges that can help researchers and practitioners to select a suitable MAS framework according to their needs.