arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

过重采样优化:大语言模型推理的测试时自校正

Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Lena Trigg, Ali Subhan, Muhammad Ali, Dean F. Hougen

arXiv 2608.05643首次发表:更新:

发表机构

University of Oklahoma; Stanford University; Universitat Pompeu Fabra; Air University(俄克拉荷马大学; 斯坦福大学; 庞培法布拉大学; 空军大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出无验证器的广度-深度优化框架,通过迭代自我批判与校正优化采样的 LLM 推理轨迹,在多个数学推理数据集及多款开放权重模型上,较基线方法实现了推理准确率的显著提升。

AI 中文摘要

测试时缩放通过增加推理计算提升大语言模型(LLM)的推理能力,但仅靠更广泛的采样会出现收益递减问题:新的推理 rollout 常重复现有答案模式,而非增加有用的推理多样性。基于验证器的选择是另一种方案,但其性能依赖外部奖励模型的校准。本文提出一种无验证器的广度-深度优化框架,利用测试时计算同时探索和改进候选解。该方法采样多个独立推理 rollout,通过迭代自我批判与自我校正优化每个 rollout,再通过多数投票聚合优化后的答案:广度保留多样化的初始尝试,深度在聚合前修复局部推理错误。在 AIME24、AIME25、AMC、OlympiadBench 和 MATH500 数据集上,该方法在多个开放权重模型上均优于贪心解码、多数投票、基于验证器的最佳 N 选择、束搜索和前瞻解码。例如,使用 Qwen2.5-1.5B 时,在 MATH500 上的准确率从最强的基于验证器的基线提升至 58.0%,在 AMC 上从 25.0% 提升至 32.5%。这些结果表明,将测试时计算用于优化采样轨迹,而非仅用于采样更多候选或依赖验证器引导的选择,会更有效。

英文摘要

Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning diversity. Verifier-based selection offers an alternative, but its performance depends on the calibration of an external reward model. We propose a verifier-free breadth--depth refinement framework that uses test-time compute to both explore and improve candidate solutions. The method samples multiple independent reasoning rollouts, refines each rollout through iterative self-critique and self-correction, and aggregates the refined answers by majority voting. Breadth preserves diverse initial attempts, while depth repairs local reasoning errors before aggregation. Across AIME24, AIME25, AMC, OlympiadBench, and MATH500, our method consistently improves over greedy decoding, majority voting, verifier-based best-of-$N$, beam search, and lookahead decoding across multiple open-weight models. For instance, with Qwen2.5-1.5B, accuracy increases from the strongest verifier-based baseline to $58.0\%$ on MATH500, and from $25.0\%$ to $32.5\%$ on AMC. These results show that test-time compute can be more effective when used to refine sampled trajectories rather than only to sample more candidates or rely on verifier-guided selection.

CommentsSubmitted to EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑