发表机构
Sun Yat-sen University(中山大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出SVR框架,通过联合判决-置信度强化学习实现自适应测试时计算,在7个数学推理基准上以更少推理轮次取得优于基线的性能,证明学习型自验证可有效控制测试时计算分配。
AI 中文摘要
扩展测试时计算可提升语言模型推理能力,但统一计算预算会在简单输入上浪费算力,而验证器引导的优化依赖外部反馈。我们提出自验证优化(Self-Verifying Refinement,SVR),这是一种无需神谕的多轮强化学习框架,学习将自验证用作计算控制策略。每一轮中,模型生成解决方案,同时输出离散的正确性判决和置信度分数;仅当判决为“正确”且置信度超过阈值时,模型保留当前答案,否则通过自身自验证继续优化。真实标签正确性仅用于构建训练奖励,优化提示或推理过程中不会向策略暴露该信息。SVR使用GRPO在固定时长轨迹上训练,奖励函数涵盖解决方案正确性、校准感知自验证及可停止的正确状态;自适应停止仅在推理阶段激活。在使用Qwen3.5-2B的7个数学推理基准上,SVR的宏观平均准确率达0.563,平均仅需2.99次推理轮次。在完整系统对比中,它优于标准GRPO、强多轮基线及固定10轮推理的神谕引导分数反馈基准,且所需轮次显著更少。这些结果表明,学习得到的自验证可作为有效的内部控制信号,用于答案保留和自适应测试时计算分配。
英文摘要
Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refinement relies on external feedback. We introduce Self-Verifying Refinement (SVR), an oracle-free multi-turn reinforcement learning framework that learns to use self-verification as a compute-control policy. At each turn, the model produces a solution together with a discrete correctness verdict and a confidence score; it retains the current answer only when the verdict is Correct and confidence exceeds a threshold, and otherwise continues refinement using its own self-verification. Ground-truth correctness is used only to construct training rewards and is never exposed to the policy through refinement prompts or required at inference. SVR is trained with GRPO on fixed-horizon trajectories using rewards that promote solution correctness, calibration-aware self-verification, and stop-ready correct states; adaptive stopping is activated only at inference. On seven mathematical reasoning benchmarks with Qwen3.5-2B, SVR achieves a macro-average accuracy of 0.563 with only 2.99 inference turns on average. In the evaluated complete-system comparison, it exceeds standard GRPO, strong multi-turn baselines, and a fixed-budget oracle-guided score-feedback reference while requiring substantially fewer turns than fixed ten-turn inference. These results demonstrate that learned self-verification can serve as an effective internal control signal for answer retention and adaptive test-time compute allocation.
Comments8 pages, 4 figures, 4 tables