arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

强化学习用于层次推理奖励:Transformer 的极小极大最优速率

Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers

Naoki Nishikawa, Taiji Suzuki

arXiv 2610.08561首次发表:更新:

发表机构

The University of Tokyo; RIKEN AIP(东京大学; 理化学研究所革新智能统合研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将推理奖励建模为层次函数,证明基于 Transformer 的演员-评论家算法在 RL 后训练中实现极小极大最优速率,优于离线采样。

AI 中文摘要

强化学习(RL)已成为对推理任务中的语言模型进行后训练的标准工具,其中策略在探索响应空间的同时根据奖励反馈进行更新。尽管其经验上取得了成功,但 RL 后训练的理论理解仍然有限,特别是关于为何在线策略探索结合神经奖励模型是有效的。在本文中,我们通过将奖励建模为响应空间上的层次函数来解决这个问题:奖励由无限多个局部组件组成,每个组件只有在前面组件被解决后才变得相关。我们证明了一种基于 Transformer 的自然演员-评论家算法,该算法在从当前 KL 正则化策略中采样、将 Transformer 评论家拟合到观察到的奖励以及更新策略之间交替进行,在查询预算和正则化强度方面(在对数因子范围内)实现了极小极大最优速率,并且对于固定数量的提示也是极小极大最优的。相比之下,我们证明从固定参考分布中采样(如在离线奖励建模中)可能将遗憾衰减限制为对数速率。这些结果表明,在线策略探索逐步聚焦于奖励集中的区域,并量化了其对 RL 后训练的益处。

英文摘要

Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses. Despite its empirical success, theoretical understanding of RL post-training remains limited, in particular of why on-policy exploration combined with a neural reward model is effective. In this paper, we address this question by modeling the reward as a hierarchical function on the response space: the reward consists of infinitely many local components, each of which becomes relevant only after the preceding ones have been resolved. We show that a natural Transformer-based actor--critic algorithm, which alternates between sampling from the current KL-regularized policy, fitting a Transformer critic to the observed rewards, and updating the policy, achieves the minimax optimal rates in the query budget and in the regularization strength up to logarithmic factors, and is minimax optimal for a fixed number of prompts. In contrast, we prove that sampling from the fixed reference distribution, as in offline reward modeling, can limit regret decay to a logarithmic rate. These results show that on-policy exploration progressively zooms in on the region where the reward is concentrated, and quantify its benefit for RL post-training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑