arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36552cs.LGcs.AI

SERA:用于最大似然强化学习的尺度均衡回放分配

SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning

Zihao Chen, Fanxiang Xiong, Hongran Ren, Xuefeng Bai, Zhongxiang Dai, Kehai Chen, Zhiguo Zhang, Zhiyong Wang, Yu Cheng

首次发表
浏览论文内容

中文总结 AI 辅助

SERA通过重新分配固定回放预算以均衡有限回放尺度因子,解决MaxRL中低成功提示梯度衰减问题,提升梯度对齐与解覆盖率。

中文摘要 AI 辅助

最大似然强化学习(MaxRL)以提示级别的对数成功率为目标,在推理任务上表现出强大的性能。然而,在有限回放预算下,MaxRL使用的估计器会以取决于每个提示的成功概率和回放次数的因子来衰减其似然梯度。在均匀回放分配下,常见的回放次数无法补偿与成功相关的衰减,导致低成功提示的衰减更严重,从而扭曲了它们对预期总梯度的相对贡献。我们提出了SERA(尺度均衡回放分配),它重新分配固定的回放预算,以近似均衡这些有限回放尺度因子。基于我们对有限回放如何扭曲提示级似然梯度的理论分析,我们将分配问题表述为固定预算的最大-最小问题,推导出其连续松弛的水线解,并引入多重性校正以消除由异构回放次数引起的额外提示加权。实验表明,在受控的ImageNet设置中,SERA与精确似然梯度的对齐更强,并且在匹配的训练回放预算下,在迷宫导航和数学推理任务上,其多样本解覆盖率优于MaxRL。

英文摘要

Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Under uniform rollout allocation, the common rollout count fails to compensate for success-dependent attenuation, leaving low-success prompts more strongly attenuated and distorting their relative contributions to the expected aggregate gradient. We introduce SERA (Scale-Equalized Rollout Allocation), which redistributes a fixed rollout budget to approximately equalize these finite-rollout scaling factors. Building on our theoretical analysis of how finite rollouts distort prompt-wise likelihood gradients, we formulate the allocation as a fixed-budget max--min problem, derive a waterline solution to its continuous relaxation, and introduce a multiplicity correction to remove the additional prompt weighting induced by heterogeneous rollout counts. Experiments show stronger alignment with exact likelihood gradients in a controlled ImageNet setting and improved multi-sample solution coverage over MaxRL on maze navigation and mathematical reasoning under matched training rollout budgets.

发表机构

  • Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑