发表机构
Zuse Institute Berlin; TU Berlin; Weizenbaum Institute Berlin(柏林Zuse研究所; 柏林工业大学; 柏林魏茨曼研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究对照评估了Self-Refine等三种LLM编排方法与基线方法在5种骨干模型、3个领域的表现,发现编排收益依赖模型,需权衡准确率提升与额外推理成本。
AI 中文摘要
人们通常认为,LLM编排通过分配额外的推理时间计算来提升推理能力,但其收益可能并不值得付出成本。现有对比研究也常常忽略优化工作量的差异,导致难以单独凸显编排本身的价值。我们针对Self-Refine、Best-of-$N$、Debate三种编排方法,以及仅任务处理、思维链(CoT)单调用基线方法,在5种LLM骨干模型和3个领域(竞赛编程、国际象棋谜题、数学)开展对照评估。为保证可比性,我们在相同优化预算下用GEPA优化每种方法,并在同一难度分层的基准项目上评估所有方法。编排带来适度但依赖基准的收益:在每个基准内对骨干模型取平均,最大提升幅度较优化后的CoT推理为4.6个百分点,较仅任务处理推理为4.5个百分点,但需要约2至4倍于仅任务处理推理的平均总token数。人工标注的难度与三个基准中更低的绝对准确率相关,但基准内分析未显示编排效果随任务难度提升。相反,探索性混合效应分析揭示,三种基准中编排方法与骨干模型间存在强交互作用,表明编排有效性在很大程度上依赖底层模型。我们的结果表明,编排决策应针对特定模型,并权衡适度准确率提升是否值得额外推理成本;更广泛而言,LLM编排的评估应控制优化工作量,并报告特定模型的准确率-成本权衡,而非将额外推理时间结构视为普遍有益。
英文摘要
LLM orchestration is often assumed to improve reasoning by allocating additional inference-time computation, yet its gains may not justify its cost. Existing comparisons also frequently overlook differences in optimization effort, making it difficult to isolate the value of orchestration itself. We conduct a controlled evaluation of Self-Refine, Best-of-$N$, and Debate against task-only and chain-of-thought (CoT) single-call baselines across five LLM backbones and three domains: competitive programming, chess puzzles, and mathematics. For comparability, we optimize each method with GEPA under the same optimization budget and evaluate all methods on the same difficulty-stratified benchmark items. Orchestration yields moderate but benchmark-dependent gains: averaged across backbones within each benchmark, the largest improvement is 4.6 percentage points over optimized CoT inference and 4.5 points over task-only inference, while requiring approximately 2 to 4 times the mean total tokens of task-only inference. Human-derived difficulty is associated with lower absolute accuracy in all three benchmarks, but within-benchmark analyses do not indicate that orchestration effects increase with task difficulty. By contrast, exploratory mixed-effects analyses reveal strong interactions between orchestration method and backbone model across all three benchmarks, showing that orchestration effectiveness depends substantially on the underlying model. Our results suggest that orchestration decisions should be model-specific and account for whether moderate accuracy gains justify the additional inference cost. More broadly, evaluations of LLM orchestrations should control optimization effort and report model-specific accuracy--cost trade-offs rather than treating additional inference-time structure as uniformly beneficial.