arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17306cs.MAcs.AI

模型越多,问题越多:设计多智能体系统时如何最优选择模型池

Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems

Sara Vera Marjanović, Jiacheng Xu, Aleksandr Laptev, Grigor Nalbandyan, Erik Arakelyan, Evelina Bakhaturina

首次发表
浏览论文内容

中文总结 AI 辅助

该研究系统评估了多智能体系统中8种模型选择策略,发现扩大候选池常降低性能,而单一模型家族内选择效果最佳,强调模型选择是关键设计决策。

中文摘要 AI 辅助

多智能体系统(MAS)结合多个模型的输出以解决复杂的推理任务。然而,尽管开源模型数量快速增长,关于如何从这一庞大模型池中选取最优候选模型的研究却十分有限。我们系统评估了8种模型选择策略(包括模型规模、准确率和答案多样性),这些策略应用于生成前(路由)和生成后(多数投票、LLM作为裁判)的MAS架构,并在具有挑战性的科学基准上进行了测试。我们的研究结果显示,理论上的最优潜力与实际性能之间存在显著差距:扩大候选池规模往往会使性能降至低于表现最佳的基础模型。我们发现,在单一模型家族内进行候选选择是相对于独立模型获得最佳相对性能的策略。这些结果表明,向异构MAS中添加任意模型可能引入系统不稳定性,凸显了模型选择作为多智能体系统设计中的关键决策。

英文摘要

Multi-agent Systems (MAS) combine multiple model outputs to solve complex reasoning tasks. However, despite rapid growth of available open-source models, there is limited research on how to select optimal model candidates out of this massive pool. We systematically evaluate 8 model selection strategies (including model size, accuracy and answer diversity) across before-generation (routing) and after-generation (majority-voting, LLM-as-a-judge) MAS architectures on challenging scientific benchmarks. Our findings show a significant gap between theoretical oracle potential and actual performance: Expanding candidate pool sizes often degrades performance below that of the top performing base-model. We find that candidate selection within a single model family is the strategy that yields the best relative performance over a standalone model. These results demonstrate that adding arbitrary models to a heterogeneous MAS can introduce system instability, highlighting model selection as a critical design choice for multi-agent systems.

发表机构

  • NVIDIA(英伟达)
  • University of Copenhagen(哥本哈根大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑