arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SLPO:通过代理策略扩展潜在推理

SLPO: Scaling Latent Reasoning via a Surrogate Policy

Runyang You, Zhiyuan Liu, Yongqi Li, Wenjie Li

arXiv 2607.19691首次发表:更新:

发表机构

The Hong Kong Polytechnic University; Sichuan University(香港理工大学; 四川大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何让潜在推理器实现测试时扩展,引入代理潜在策略优化(SLPO),通过经验代理策略密度和正确性监督停止头,将结果奖励强化学习引入自回归潜在推理器,在连续和软思维设置中提升了性能。

AI 中文摘要

具有可验证奖励的强化学习已成为在显式思维链推理器中实现测试时扩展的主要方法。然而,这种扩展路径计算成本高昂,因为每个中间步骤都必须解码为语言令牌。潜在推理则将中间计算作为连续向量进行,并且在更短的时间范围内已经匹配或超过了显式思维链。尽管有此前景,但潜在推理器在很大程度上仍受模仿限制,而显式思维链已通过结果奖励强化学习超越了模仿。潜在轨迹在固定思维预算下缺乏易于处理的每步似然性和自适应停止接口,因此结果奖励无法引发潜在测试时的扩展。我们引入代理潜在策略优化(SLPO),将结果奖励强化学习引入自回归潜在推理器:一种用于轨迹级信用分配的潜在转换的经验代理策略密度,以及一个正确性监督停止头,结果奖励优化将其细化为可变时间范围策略。在连续和软思维设置中,SLPO在并行采样下提高了Pass@$k$,并将更长的潜在计算分配给具有更高确定性准确性的更难实例。

英文摘要

Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since intermediate reasoning must be externalized as natural-language tokens. Latent reasoning instead carries intermediate computation as continuous vectors and already matches or surpasses explicit CoT at far shorter horizons. Despite this promise, latent reasoners remain largely imitation-bound, while explicit CoT has already moved past imitation via outcome-reward RL. Latent trajectories lack a tractable per-step likelihood and an adaptive stopping interface under fixed thinking budgets, so outcome rewards cannot elicit latent test-time scaling. We introduce Surrogate Latent Policy Optimization (SLPO) to bring outcome-reward RL to autoregressive latent reasoners: a differentiable surrogate policy interface over latent transitions for trajectory-level credit assignment, and a correctness-supervised stopping head that outcome-reward optimization refines into a variable-horizon policy. Across two continuous latent reasoners, two backbones, and three held-out benchmarks, SLPO improves Pass@8 and Pass@16 in all 12 backbone--dataset settings, with gains of up to 12.07 percentage points. SLPO further transfers to soft-token inference and learns difficulty-adaptive computation, allocating longer latent trajectories to harder instances. Project Page: https://modalitydance.github.io/SLPO/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑