在相同推理成本下,多智能体结构无法胜过单个冻结智能体
At Equal Inference Cost, Multi-Agent Structure Does Not Beat a Single Frozen Agent
浏览论文内容
中文总结 AI 辅助
该研究在相同推理成本下对比多智能体团队与单个冻结智能体的性能,提出MA-Evolve方法,发现多智能体结构无明显优势,仅执行器贡献性能提升。
中文摘要 AI 辅助
多智能体大语言模型(LLM)流水线,如规划器-执行器-评判器团队,通常报告比单个智能体有性能提升,但这些提升通常伴随着更高的推理成本,因为团队在每个环境步骤中会进行多次模型调用。现有的自动化方法会在角色、拓扑结构和提示之间进行搜索,但通常在相等的环境部署次数下将团队与单个智能体进行比较,这给了团队额外的计算资源。我们相反地固定语言模型调用的总数,询问在相同预算下演化多智能体团队是否仍然胜过演化单个智能体。我们引入MA-Evolve,它将规划器-执行器-评判器团队表示为三个可演化的角色提示,并通过在共享的冻结7B主干上按角色进行坐标上升来优化它们。在ALFWorld上,演化单个执行器显著优于未演化的智能体,而完整团队达到最高均值,但与单个智能体相比没有统计学上的显著差异:0.769对0.754,p=0.80,尽管使用了1.8倍的评估调用。留一法分析显示,实现的价值完全来自执行器;规划器和评判器演化成空或低影响的提示,很少改变执行器的动作。在2-3倍的免费计算下,团队仅与单个智能体匹配,而在WebShop上演化无效,团队趋势更差。在相同推理成本下,多智能体结构增加了成本却没有明显的益处。
英文摘要
Multi-agent LLM pipelines, such as Planner-Executor-Critic teams, often report gains over single agents, but these gains usually come with higher inference cost because the team makes multiple model calls per environment step. Existing automated methods search over roles, topologies, and prompts, but typically compare teams against single agents at equal environment rollouts, giving the team extra compute. We instead fix the total number of language-model calls and ask whether evolving a multi-agent team still beats evolving a single agent under the same budget. We introduce MA-Evolve, which represents a Planner-Executor-Critic team as three evolvable role prompts and optimizes them by per-role coordinate ascent over a shared frozen 7B backbone. On ALFWorld, evolving a single executor significantly improves over the unevolved agent, while the full team achieves the highest mean but is not statistically better than the single agent: 0.769 versus 0.754, p = 0.80, despite using 1.8 times more evaluation calls. Leave-one-in analysis shows that the realized value comes entirely from the executor; the planner and critic evolve to empty or low-impact prompts and rarely change the executor's action. With 2-3 times free compute, the team only matches the single agent, and on WebShop evolution is null while the team trends worse. Under equal inference cost, multi-agent structure adds cost without clear benefit.
发表机构
- Trinity College Dublin(都柏林三一学院)
- University College Dublin(都柏林大学)
- Dublin City University(都柏林城市大学)
机构由 AI 辅助整理,请以论文原文为准。