多智能体在析取性与补偿性任务中的规模扩展
Multi-agent Scaling Across Disjunctive and Compensatory Tasks
AI总结:
本研究引入Steiner分类法分析多智能体LLM扩展,发现任务结构决定扩展效果:析取性任务中多数投票增益有限,补偿性任务中平均化收益微小。
AI中文摘要:
多智能体LLM系统通常预期随着团队规模的增大而性能提升,然而其扩展行为可能取决于任务结构。我们的核心贡献是引入Steiner的群体任务分类法作为分析多智能体LLM扩展的框架,并将分析聚焦于析取性任务和补偿性任务。我们将独立采样的智能体建模为在给定项目条件下条件独立,由此得出其大规模团队极限:多数投票收敛于模型的众数答案,而平均化收敛于模型的项目级偏差。在选定的代表性基准、13个开放权重模型以及多达30个智能体的团队中,我们发现定性不同的扩展行为。在析取性任务上,至少一个智能体正确的概率随团队规模增长5-20个百分点,但对直接作答的智能体进行多数投票几乎无法实现这一潜力,因为模型平均预测偏差在0.5个百分点以内。多轮修订显著提高准确率,但增益与拥有1个同伴时几乎相同,即使有29个同伴也是如此。相比之下,在费米估计上扩展几乎无益,尽管其天然适合聚合:模型样本间共享的项目级偏差约占平方误差的87%,因此平均化仅减少约6%的误差。组合模型家族在费米估计上有帮助,但在析取性任务上未超过最强成员。这些结果表明,任务结构连同组合成员输出的机制,是团队扩展的基本决定因素。
英文摘要:
Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory tasks. We model independently sampled agents as conditionally independent given the item, which yields their large-team limits: plurality voting converges to the model's modal answer, and averaging converges to the model's item-level bias. Across selected representative benchmarks, 13 open-weight models, and teams of up to 30 agents, we find qualitatively different scaling behavior. On disjunctive tasks, the probability that at least one agent is correct grows by 5-20 points with team size, but plurality voting over agents that answer directly realises almost none of this potential, as the model predicts to within 0.5 points on average. Multi-round revision raises accuracy considerably, yet the gain is nearly the same with one peer as with 29. In contrast, scaling provides little benefit on Fermi estimation, despite its natural suitability for aggregation: item-level biases shared across the samples of a model account for about 87% of the squared error, so averaging reduces error by only about 6%. Combining model families helps on Fermi estimation but does not surpass the strongest member on disjunctive tasks. These results show that task structure, together with the mechanism combining member outputs, is a fundamental determinant of team scaling.