arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19499cs.LGcs.DCcs.PF

样本数量并不足够:候选生成策略决定LLM测试时扩展的能量与性能

Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling

Mobina Kashaniyan, Ali Jannesari

首次发表
浏览论文内容

中文总结 AI 辅助

研究发现,LLM测试时扩展中,候选数量相同但生成调度不同(批量vs串行)会导致显著的能量和延迟差异,批量生成更高效,评估应报告调度和GPU指标。

中文摘要 AI 辅助

测试时扩展可以通过生成并组合多个候选响应来提升大语言模型的推理能力。在基于采样的方法中,推理预算通常由生成的候选数量N来描述。然而,N只告诉我们生成了多少个候选,而没有说明它们是如何执行的。相同的候选预算可以在一次批量生成调用中产生,也可以拆分为多次批量大小更小的顺序调用。我们首先使用Phi-3-mini和Qwen2.5-1.5B在500个GSM8K提示上研究增加N对推理准确率的影响。正如预期,将N从1增加到8,Phi-3-mini的准确率提高了8.4个百分点,Qwen2.5-1.5B提高了18.4个百分点。然而,仅凭准确率并不能显示使用更大候选预算的系统成本。因此,我们固定N=8,比较四种生成调度:1x8、2x4、4x2和8x1,其中axb表示a次生成调用,每次调用生成b个候选。在保持总候选数固定的情况下,我们测量了延迟、吞吐量、GPU小时数和GPU设备总能耗。在A100 GPU上,八次串行调用使用的GPU设备总能耗是单次批量调用(含八个候选)的4.64-4.86倍,P95延迟是其5.77-6.12倍。这一模式在每种模型三个独立调度的A100节点以及短输出SciQ/V100实验中均一致出现。这些结果表明,仅凭候选数量不足以描述多候选测试时扩展的系统成本。当候选相互独立且内存允许时,更少的生成调用配合更大的批量大小更为高效。因此,评估不仅应报告候选数量和准确率,还应报告生成调度和GPU级系统指标。

英文摘要

Test-time scaling can improve large language model reasoning by generating and combining multiple candidate responses. In sampling-based methods, the inference budget is often described by the number of generated candidates, N. However, N tells us how many candidates are generated, not how they are executed. The same candidate budget can be produced in one batched generation call or split across several sequential calls with smaller batch sizes. We first study the effect of increasing N on reasoning accuracy using Phi-3-mini and Qwen2.5-1.5B on 500 GSM8K prompts. As expected, increasing N from 1 to 8 improves accuracy by 8.4 percentage points for Phi-3-mini and 18.4 points for Qwen2.5-1.5B. However, accuracy alone does not show the systems cost of using a larger candidate budget. We therefore fix N = 8 and compare four generation schedules: 1x8, 2x4, 4x2, and 8x1, where axb denotes a generation calls with b candidates per call. We measure latency, throughput, GPU-hours, and gross GPU-device energy while keeping the total candidate count fixed. On A100 GPUs, eight serial calls use 4.64-4.86x as much gross GPU-device energy and have 5.77-6.12x the P95 latency of one batched call with eight candidates. The same pattern appears across three independently scheduled A100 nodes per model and in short-output SciQ/V100 experiments. These results show that candidate count alone is not enough to describe the systems cost of multi-candidate test-time scaling. When candidates are independent and memory allows it, fewer generation calls with larger batch sizes are more efficient. Evaluations should therefore report not only candidate count and accuracy, but also generation schedule and GPU-level systems metrics.

发表机构

  • Iowa State University(爱荷华州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑