扩展大语言模型驱动的多智能体系统:设计原则与架构可扩展性分析
Scaling LLM-Driven Multi-Agent Systems: Design Principles and Architectural Scalability Analysis
AI总结:
该研究提炼出可扩展多智能体系统的四项设计原则,构建参考架构并在基准任务上评估,发现LLM超能力阈值时扩展可提升准确率但成本近似线性增长,性能在中等复杂度达峰,一致性是核心挑战。
AI中文摘要:
基于大语言模型(LLM)的多智能体系统(MAS)有潜力通过专用智能体的协同集合实现集体智能,并扩展以解决高度复杂的任务。然而,尽管其具有理论潜力,架构设计空间在很大程度上仍未系统化,缺乏广泛确立的设计原则;此外,此类系统的可扩展性特征目前仅被部分理解。本文做出两项贡献:首先,通过对现有研究的结构化分析,提炼出可扩展MAS架构的四项设计原则:简洁性、弹性反馈、带可选循环的顺序工作流,以及基于摘要的通信。我们将这些原则应用于参考架构,其拓扑结构被形式化为受约束的有向工作流图,并使用两种能力不同的LLM,在基于终端的系统工程任务标准化基准上,评估了四种复杂度递增的配置。研究发现,当底层LLM超过最低能力阈值时,扩展会带来可测量的准确率提升,且成本增长近似线性;性能在中等复杂度时达到峰值,随后因超时和评估限制而下降。此外,在所有扩展级别中,持续的一致性问题成为核心挑战。这些结果为从业者提供了具体的设计指导,并强调一致性和评估标准化是未来研究的关键目标。
英文摘要:
LLM-based multi-agent systems have the potential to enable collective intelligence and scale toward solving highly complex tasks through coordinated ensembles of specialized agents. However, despite their theoretical potential, the architectural design space remains largely non-systematized and lacks broadly established design principles. Furthermore, the scalability characteristics of such systems are only partially understood so far. This paper makes two contributions. We first distill four design principles for scalable MAS architectures from a structured analysis of prior work: simplicity, elastic feedback, sequential workflows with optional loops, and summary-based communication. We operationalize these principles in a reference architecture whose topology is formalized as a constrained directed workflow graph, and we evaluate four configurations of increasing complexity on a standardized benchmark of terminal-based system engineering tasks using two LLMs of differing capability. Our findings show that scaling yields measurable accuracy improvements with approximately linear cost growth, but only when the underlying LLM exceeds a minimum capability threshold. Performance peaks at intermediate complexity, then degrades due to timeouts and evaluation limitations. In addition, persistent consistency issues emerge as a central challenge across all scaling levels. These results provide concrete design guidance for practitioners and highlight consistency and evaluation standardization as key targets for future research.