arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

小模型团队扩展在编排架构上的精确生成-变换分解

An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration Architectures

Blaz Bertalanic, Carolina Fortuna

arXiv 2609.36104首次发表:更新:

发表机构

Jožef Stefan Institute(约瑟夫·斯特凡研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究团队扩展在不同编排架构下的效果,发现其回报高度依赖任务,Proposer-Critic在算术任务上表现最佳,并通过精确分解揭示了覆盖与变换的作用。

AI 中文摘要

将单个LLM代理替换为协作团队可以提高准确性,但团队扩展是否有帮助,以及应该扩展哪种架构,目前尚不清楚。我们横跨五种指令微调的7-9B模型、五个短答案基准和一个可执行代码基准(最多30次调用),对八种代理编排架构进行了全面扫描,发现团队扩展的回报高度依赖于任务:从三次到三十次调用,在两个算术应用题基准(GSM8K、GSMHard)上,准确性最多提升17个百分点,但在ARC、GPQA和MMLU上,每种架构的提升最多不超过4个百分点,这一差异通常被任务平均数字所掩盖。Proposer-Critic架构捕捉了算术上的增益,扩展最为陡峭,并且在总体上,在最大预算下超越了所有其他架构(项目聚类区间排除零),尽管它在其他任务中排名最弱,且没有一种架构在所有任务中获胜。我们用一种精确的生成-变换分解来解释这些轨迹。将任何工作流划分为提案覆盖和下游变换,任何准确性的变化都精确地分解为广泛的覆盖红利和密集的变换变化。该分解诊断了每个任务:算术提供了覆盖空间,评论家引导的变换可以转化这些空间,而多项选择基准要么在覆盖上饱和,要么未能转化覆盖,在开放式代码生成恢复几乎消失,因此准确性跟随覆盖。在相同的调用预算下,令牌成本仍然变化2.1倍。因此,额外的调用创造了候选机会,只有某些架构在某些任务上能转化这些机会。团队扩展是一个任务和架构特定的赌注,而不是一个统一的杠杆。

英文摘要

Replacing one LLM agent with a collaborating team can raise accuracy, but whether scaling the team helps, and which architecture to scale, is unclear. Sweeping eight agent orchestration architectures across five instruction-tuned 7-9B models, five short-answer benchmarks, and an executable-code benchmark up to 30 calls, we find that the returns to team scaling are sharply task-dependent: from three to thirty calls accuracy rises by up to 17 points on the two arithmetic word-problem benchmarks (GSM8K, GSMHard) but by at most four on ARC, GPQA, and MMLU, for every architecture, a split the usual task-averaged number conceals. Proposer-Critic captures the arithmetic gains, scaling steepest and, in aggregate, surpassing every other architecture at the largest budget (item-clustered intervals exclude zero), though it ranks among the weakest elsewhere, and no architecture wins across tasks. We explain these trajectories with an exact generate-transform decomposition. Partitioning any workflow into proposal coverage and a downstream transform, any accuracy change splits exactly into an extensive coverage dividend and an intensive transformation change. The decomposition diagnoses each task: arithmetic offers coverage headroom that a critic-guided transform converts, whereas the multiple-choice benchmarks either saturate in coverage or fail to convert it, and on open-ended code generative recovery nearly vanishes so accuracy tracks coverage. At equal call budgets token cost still varies 2.1x. Extra calls therefore create candidate opportunity that only some architectures, on some tasks, convert. Team scaling is a task- and architecture-specific bet, not a uniform lever.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑