发表机构
Racah Institute of Physics; The Hebrew University of Jerusalem(拉卡物理研究所; 耶路撒冷希伯来大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过自旋玻璃理论证明,在迭代乘法等算法任务上,RLVR的优化景观是良性的,实际困难源于扩散障碍和梯度估计误差,可通过熵正则化缓解,且Transformer能成功学习此类任务。
AI 中文摘要
尽管具有可验证奖励的强化学习(RLVR)的重要性,但其能在多大程度上学习新的推理能力仍存在争议。在此,我们研究了RLVR在算法任务(如迭代群和拟群乘法)上的优化景观。为此,我们将近视表格策略上的熵正则化RLVR映射到确定性策略上的基于能量(自旋玻璃)模型。该映射给出了RLVR所能达到的上界,并使我们能够严格刻画这种表格设置下的景观。我们在理论和实验上均表明,对于一类广泛的模型和具有不相关输入的任务,该景观是良性的,不包含可能困住RLVR训练的局部极小值。相反,这些任务的实际困难似乎至少部分源于诸如扩散障碍和在穿越景观时的梯度估计误差等问题。这些是可能阻止找到解决方案的真正障碍,但它们与景观本身崎岖不同。我们表明,这些障碍通常可以通过选择熵正则化器来缓解。与该理论一致,我们发现一个从头训练、仅使用最后令牌奖励的Transformer成功学习了迭代非阿贝尔群乘法的算法思维链。
英文摘要
Despite the importance of reinforcement learning with verifiable rewards (RLVR), the extent to which it can learn new reasoning capabilities remains debated. Here we study the optimization landscape of RLVR on algorithmic tasks, such as iterated group and quasigroup multiplication. To this end, we map entropy-regularized RLVR over myopic tabular policies onto an energy-based (spin-glass) model over deterministic policies. This mapping upper-bounds what RLVR can achieve, and lets us rigorously characterize the landscape in this tabular setting. We show, both theoretically and experimentally, that for a wide class of models and tasks with uncorrelated inputs, this landscape is benign, containing no local minima that could trap RLVR training. Rather, the practical difficulty of these tasks appears to stem, at least in part, from issues such as diffusive barriers and gradient-estimation error in traversing the landscape. These are genuine obstacles that can prevent a solution from being found, but they are distinct from the landscape itself being rugged. We show that these obstacles can often be mitigated through the choice of entropy regulator. Consistent with this theory, we find that a transformer trained from scratch, using only last-token rewards, successfully learns an algorithmic chain of thought for iterated non-Abelian group multiplications.