发表机构
Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大型语言模型复杂任务性能提升问题,提出通过错误定位进行测试时缩放(TTEL)算法,利用反馈定位错误步骤,截断轨迹并分支新一代,重用有效前缀,在多个基准测试中优于竞争基线。
AI 中文摘要
扩大推理时间计算已成为提高大型语言模型在复杂推理和编程任务上性能的可靠方法。然而,诸如独立采样和顺序多轮细化等标准方法在没有令牌级信用分配的情况下运行,导致计算效率低下。本文介绍了通过错误定位进行测试时缩放(TTEL),一种利用固定或环境反馈进行令牌级错误定位的推理算法。通过将有信息反馈下的条件概率与空上下文基线进行比较,TTEL隔离错误发生的步骤,然后截断轨迹并分支新一代,最大限度地重用有效前缀。大量评估表明,TTEL在顺序推理领域建立了严格占优的帕累托前沿。在LiveCodeBench上使用Qwen3-8B时,TTEL在生成大约独立采样一半令牌数量(360.4k对735.0k)的情况下,达到了71.0%的pass@64。在数学基准测试AIME-2025和HMMT-2025中,TTEL在Qwen3-8B和Qwen3-4B-Thinking-2507上均明显优于竞争的测试时基线。
英文摘要
Scaling inference-time computation has emerged as a reliable method to improve the performance of large language models on complex reasoning and programming tasks. However, standard approaches such as independent sampling and sequential multi-turn refinement operate without token-level credit assignment, resulting in computational inefficiency, since valid reasoning prefixes are frequently discarded. In this work, we introduce Test-Time Scaling via Error Localization (TTEL), an inference-time algorithm that utilizes fixed or environment feedback to perform token-level error localization. By comparing conditional probabilities under informed feedback against a null-context baseline, TTEL isolates the step at which an error occurred. The algorithm then truncates the trajectory and branches a new generation, maximally reusing the valid prefix. Extensive evaluations demonstrate that TTEL establishes strictly dominating Pareto frontiers across sequential reasoning domains, measured by pass-at-k vs. generated-token cost. With Qwen3-8B on LiveCodeBench, TTEL attains a pass@64 of 71.0% while generating approximately half as many tokens as independent sampling (360.4k vs. 735.0k). Generalizing to math benchmarks AIME-2025 and HMMT-2025, TTEL cleanly outperforms competing test-time baselines across both Qwen3-8B and Qwen3-4B-Thinking-2507.
Comments10 pages, 8 figures (With appendix: 27 pages, 11 figures)