发表机构
University of Notre Dame; Amazon, Inc(圣母大学; 亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有大语言模型后训练中展开分配的局限,本文提出RAIL框架,将干预选择建模为在线上下文博弈问题,可在有限展开预算下提升性能,为生成高质量展开提供原则性方法。
AI 中文摘要
无批评者的基于群体的强化学习已成为大语言模型后训练的可扩展方法,但现有多数方法为每个任务和轨迹状态分配相同数量的展开,而部分展开能提供更有用的学习信号。近期研究开始将展开生成视为自适应决策,但仍存在两个关键局限:其一,干预策略常基于固定启发式,无法随训练中策略变化调整;其二,这些方法通常仅决定生成多少展开,未明确控制干预的位置与方式。为解决这些局限,本文提出可恢复性感知干预学习(RAIL),这是一种训练时框架,基于每次干预产生的改进学习如何生成展开。RAIL将干预选择建模为在线上下文博弈问题,通过影子到真实程序收集的干预轨迹训练可恢复性控制器,使控制器能在底层策略演化时持续学习。本文从有效性、适应性、表达性和效率四个方面评估RAIL,在多种设置下,RAIL在有限展开预算下始终提升性能。这些结果表明,可恢复性感知干预提供了一种生成更具信息性、更少冗余的展开的原则性方法,在后训练期间产生更强的学习信号。
英文摘要
Critic-free group-based reinforcement learning has become a scalable approach for post-training large language models. However, most existing methods allocate the same number of rollouts to every task and trajectory state, even though some rollouts provide much more useful learning signals than others. Recent work has started to treat rollout generation as an adaptive decision, but two important limitations remain. First, intervention strategies are often based on fixed heuristics and therefore cannot adjust as the policy changes during training. Second, these methods usually decide only how many rollouts to generate, without explicitly controlling where and how to intervene. To address these limitations, we propose Recoverability-Aware Intervention Learning (RAIL), a training-time framework that learns how to generate rollouts based on the improvement produced by each intervention. RAIL models intervention selection as an online contextual-bandit problem and trains a recoverability controller using intervention traces collected through a shadow-to-live procedure. This allows the controller to keep learning while the underlying policy evolves. We evaluate RAIL in terms of effectiveness, adaptivity, expressiveness, and efficiency. Across multiple settings, RAIL consistently improves performance under limited rollout budgets. These results show that recoverability-aware intervention provides a principled way to generate more informative and less redundant rollouts, leading to stronger learning signals during post-training.