arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38274cs.CLcs.LGcs.MA

哪些模型协同工作效果好?用于LLM团队选择的异质性度量

Which Models Work Well Together? Measuring Heterogeneity for LLM Team Selection

  • Fudan Institute on Networking Systems of AI(复旦大学人工智能网络系统研究院)

机构由 AI 辅助整理,请以论文原文为准。

Liangyu Teng, Hengsong Liu, Juncen Guo, Jingyu Zhang, Yang Liu, Jing Liu, Liang Song

AI总结:

提出异质性驱动的LLM团队选择框架,通过错误去相关和预测差异度量互补性,结合贪心搜索优化团队组合,实验证明优于仅质量基线。

AI中文摘要:

LLM团队的性能上限不仅受限于单个模型的能力,还受限于成员间的错误共振和预测差异。尽管异质性团队在实践中常被观察到有效,现有方法缺乏可计算、可解释且可优化的互补性度量,导致团队组建依赖启发式方法。我们提出一个异质性驱动的团队选择框架,通过离线画像刻画个体能力,并结合两种互补信号:一种捕获错误模式去相关以减少共同失败,另一种度量预测行为差异以获取策略多样性。我们将团队选择形式化为标准化的质量-互补性组合目标,并应用高效贪心搜索从候选池中选取小型团队。跨多个基准的实验表明,在受控候选池和团队规模下,我们的框架始终优于仅基于质量的基线,为多LLM系统建立了可复用的选择原则。

英文摘要:

The performance ceiling of an LLM team is constrained not only by individual model capabilities, but also by inter-member error resonance and predictive differences. Although heterogeneous teaming is often observed to be effective in practice, existing approaches lack complementarity metrics that are computable, interpretable, and optimizable, leaving team composition to rely on heuristics. We propose a heterogeneity-driven team selection framework that performs offline profiling to characterize individual capability along with two complementary signals: one captures decorrelation in error patterns to reduce co-failures, while the other measures divergence in predictive behavior to capture strategy diversity. We formulate team selection as a standardized quality--complementarity combinatorial objective and apply an efficient greedy search to select a small team from a candidate pool. Experiments across multiple benchmarks demonstrate that our framework consistently outperforms quality-only baselines under controlled candidate pools and team sizes, establishing reusable selection principles for multi-LLM systems.

↑