希尔采样用于测试时扩展:一种比重复采样、进化和训练更简单且更优的替代方案
Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training
- Oracle(甲骨文公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出希尔采样,通过重复采样编辑并保留最优解,在圆填充等问题上超越进化方法,表明简单测试时采样优于复杂机制。
AI中文摘要:
大型语言模型(LLMs)可以通过在测试时花费额外的计算来改进可验证的科学和算法问题的解决方案。近期系统通过日益复杂的进化搜索框架或在测试时训练期间更新模型参数,取得了强劲的结果。我们探究这些机制中有多少是必要的。我们引入了希尔采样(Hill Sampling),这是一种简单的程序,它从冻结的LLM中重复采样候选程序编辑,保留目前找到的最佳程序,并让所有后续采样都以该程序为条件。我们使用三个开放权重模型,在圆填充、集合的和/差以及Erdos最小重叠问题上评估了该方法。希尔采样在圆填充问题上创下了已发表方法中的新最优水平,在Erdos最小重叠问题上优于AlphaEvolve参考方法,并在有限集合的和与差上取得了强劲结果。圆填充和Erdos结果仅需在八块NVIDIA H100 GPU上花费数小时的墙钟时间。据我们所知,我们还进行了按参数数量计的最大规模的进化策略(ES)研究,该策略在测试时直接应用于LLM权重。令人惊讶的是,学习权重比将ES学习率设为零更差:在零学习率下,该方法仍然通过固定的随机扰动在权重空间中进行搜索。这些扰动有助于探索,但来自令牌采样的随机性更强,而重复采样仍然明显弱于希尔采样。这些结果表明了一种简单的测试时计算分配策略:在引入额外的复杂性(如添加档案、多样性机制、进化脚手架或测试时参数学习)之前,反复对目前找到的最佳验证解决方案进行编辑采样。
英文摘要:
Large language models (LLMs) can improve solutions to verifiable scientific and algorithmic problems by spending additional computation at test time. Recent systems achieve strong results with increasingly elaborate evolutionary search harnesses or by updating model parameters during test-time training. We ask how much of this machinery is necessary. We introduce Hill Sampling, a form of hill-climbing optimization that repeatedly samples candidate programs from a frozen LLM, retains the best program found so far, and conditions all subsequent samples on that program. We evaluate the method on circle packing, sums and differences of sets, and Erdos' minimum-overlap problem using three open-weight models. Hill Sampling sets a new state of the art on circle packing among published methods, improves over the AlphaEvolve reference on Erdos' minimum-overlap problem, and achieves strong results on sums and differences of finite sets. The circle-packing and Erdos results require only hours of wall-clock time on eight NVIDIA H100 GPUs. We compare Hill Sampling against what is, to our knowledge, the largest application by parameter count of evolution strategies (ES) to LLM weights at test time. Surprisingly, when evaluating a method by the best program it generates, we find that learning model weights via ES is worse than invoking ES with a learning rate set to zero, i.e., using random weight-space perturbations to search for better models. Moreover, repeated sampling outperforms both ES methods, and Hill Sampling is the strongest of all. These results suggest a simple test-time compute allocation strategy: repeatedly sample edits to the best verified solution found so far before introducing additional complexity, such as adding archives, diversity mechanisms, evolutionary scaffolds, or test-time parameter learning.