arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38816cs.CL

你被录用了:面向大语言模型协作的战略性模型选择

You're Hired: Strategic Model Selection for LLM Collaboration

Zongwan Cao, Ziyuan Yang, Shangbin Feng, Michael Duan, Skyler Hallinan, Bingbing Wen, Lucy Lu Wang, Yulia Tsvetkov

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出并评估9种模型选择算法以优化多LLM协作,实验表明基于能力与训练的策略优于随机或启发式方法,最高提升36.1%,并推荐在部署前采用以过滤不安全模型并泛化至新任务。

中文摘要 AI 辅助

尽管多智能体和模型协作算法因结合不同大语言模型(LLM)的优势而日益受到关注,现有系统仍受限于预定义和手工设计的模型池。在本工作中,我们研究多LLM系统中的模型选择问题。我们提出并系统评估了一个包含9种选择算法的分类体系,涵盖模型描述的多样性、基于能力的行文多样性以及基于LLM的招聘者。我们在两个候选池(包含10个和32个模型)中进行了广泛实验,部署于四种模型协作算法,并在涵盖数学、编程、问答和推理的任务上进行了评估。结果表明,成功的选择算法在各项设置中大幅优于随机或基于启发式的团队(如仅选择个体性能最高的模型),最高提升达36.1%。具体而言,基于能力和训练的选择策略缓解了选择方差并实现了最佳性能,我们建议在部署真实世界的多LLM系统之前采用这些策略。进一步分析揭示,更大的候选池对浅层选择启发式方法构成更大挑战,而基于与候选模型交互并理解模型能力的算法能够稳健地过滤掉不匹配、不安全的模型,并泛化到新颖的、分布外的任务。综合来看,我们确立了有原则且信息充分的团队选择至关重要,并提出了用于组建高效多LLM系统的强模型选择算法。

英文摘要

While multi-agent and model collaboration algorithms gain traction to combine the strengths of diverse Large Language Models (LLMs), existing systems remain bottlenecked on pre-defined and hand-crafted model pools. In this work, we investigate the problem of model selection in multi-LLM systems. We propose and systematically evaluate a taxonomy of 9 selection algorithms ranging from diversity of model descriptions, capability-aware behavioral diversity, and LLM-based recruiters. We conduct extensive experiments across two candidate pools of 10 and 32 models, deployed in four model collaboration algorithms, and evaluated across tasks spanning math, coding, QA, and reasoning. Results demonstrate that successful selection algorithms greatly outperform random or heuristics-based teams such as merely selecting the models with top individual performance, by up to 36.1% across settings. Specifically, capability- and training-based selection strategies alleviate selection variance and achieve the best performance, which we recommend to employ before deploying real-world multi-LLM systems. Further analysis reveals that larger candidate pools pose greater challenges to shallow selection heuristics, while algorithms grounded in interacting with candidate models and understanding model capability robustly filter out misaligned, unsafe models, as well as generalizing to novel, out-of-distribution tasks. Together, we establish that principled and informed team selection is critical and present strong model selection algorithms for assembling effective multi-LLM systems.

发表机构

  • University of Washington(华盛顿大学)
  • University of Southern California(南加州大学)
  • Allen Institute for AI(艾伦人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑