arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向更好的顺序测试时扩展中的探索

Towards Better Exploration in Sequential Test-Time Scaling

Joseph Rance, Fabio Pizzati, Juil Sock, Woody Bayliss, Marc Górriz Blanch, Philip Torr, Adel Bibi

arXiv 2609.39632首次发表:更新:

发表机构

University of Oxford; MBZUAI; BBC R&D(牛津大学; 穆罕默德·本·扎耶德人工智能大学; BBC研发部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对顺序测试时扩展陷入吸引子而停止改进的问题,提出模型混合干预以逃离吸引子,显著降低命中率并提升覆盖范围和准确率,倡导转向顺序方法。

AI 中文摘要

测试时扩展通过在推理时增加额外计算来提升语言模型的推理能力。然而,现有两类方法往往在长时间尺度上难以持续改进。并行方法从模型中重复采样独立答案,在模型单次尝试难以解决的问题上扩展效果不佳。相比之下,顺序方法基于先前答案来获取新思路,但迄今为止尚未被证明能达到超出并行扩展所发现的答案。首先,我们表明顺序扩展常常因过早陷入吸引子而停止改进:吸引子是一组答案,一旦进入就会阻止对不同答案的探索。在扩展方法、模型和基准的27种组合中,我们发现53.8%的顺序扩展轨迹在四次迭代内进入吸引子。其次,我们表明一种简单的模型混合干预有助于逃离吸引子。这平均将吸引子命中率降低了21.2个百分点,将解覆盖范围扩展到超出计算匹配的并行基线,并将递归自聚合的准确率提高了至少2.2个百分点。我们的结果促使将长时程测试时扩展的关注点从并行方法转向改进先前答案的顺序方法。

英文摘要

Test-time scaling improves language model reasoning by spending additional compute at inference. However, both classes of existing methods often fail to continue improving over long timescales. Parallel methods repeatedly sample independent answers from the model, scaling poorly on problems the model is unlikely to solve in a single attempt. In contrast, sequential methods build on previous answers to access new ideas, yet so far have not been shown to reach answers beyond those found by parallel scaling. First, we show that sequential scaling often stops improving because it becomes prematurely trapped in an attractor: a set of answers that prevents exploration of different answers once entered. Across 27 combinations of scaling methods, models, and benchmarks, we find that 53.8% of sequential scaling trajectories enter an attractor within four iterations. Second, we show that a simple model-mixing intervention helps escape attractors. This reduces the attractor hit rate by 21.2 percentage points on average, expands solution coverage beyond a compute-matched parallel baseline, and improves accuracy of recursive self-aggregation by at least 2.2 percentage points. Our results motivate refocusing long-horizon test-time scaling from parallel methods to sequential methods that improve previous answers.

Comments26 pages, 15 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑