发表机构
MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PoEM框架利用现有策略预测新奖励函数下的强化学习结果,无需额外RL训练,通过线性组合对数策略实现,并在文本和图像任务上验证有效。
AI 中文摘要
基础模型通过强化学习(RL)进行后训练,以最大化特定奖励,如人类对齐、正确性或指令遵循。这一后训练过程计算密集、有时不稳定,且每当奖励模型改变或需要组合多个奖励时,都必须从头开始运行。因此,我们提出一个问题:给定一个新的奖励函数,是否有可能在不实际运行RL的情况下预测RL的结果?我们通过引入PoEM框架对此给出肯定回答,该框架利用一组已在其他奖励上完成后训练的模型,来预测新奖励函数上RL的输出。首先,我们证明如果新奖励函数可以写成现有奖励的线性组合,那么对数空间中的新策略可以写成现有对数策略的线性组合。令人惊讶的是,即使在奖励并非线性相关的情况下,我们也观察到RL训练得到的对数策略通常跨越一个近似低秩的子空间。有利的是,这种组合的加权系数仅需使用样本上的奖励或基础策略输出即可估计。我们将这些观察转化为一种算法,该算法接收后训练模型和新奖励函数,无需运行任何额外的RL训练即可近似目标RL策略。我们通过涵盖文本和图像模态的合成奖励和真实奖励实验验证了我们的方法。
英文摘要
Foundation models are post-trained with reinforcement learning (RL) to maximize specific rewards, such as human alignment, correctness, or instruction following. This post-training process is computationally intensive, sometimes unstable, and has to be run from scratch every time the reward model changes or when we want to combine multiple rewards. We hence ask: given a new reward function, is it possible to predict the RL outcomes without actually running RL on it? We answer this in the affirmative by introducing PoEM, a framework to predict the outputs of RL on a new reward function using a set of models already post-trained on other rewards. First, we show that if the new reward function can be written as a linear combination of existing ones, then the new policy in log-space can be written as a linear combination of the existing log-policies. Surprisingly, even in cases where the rewards are not linearly connected, we observe that often log-policies from RL training span an approximately low-rank subspace across rewards. To our benefit, the weighting coefficients for this combination can be estimated using only the reward or basis policy outputs on the samples. We turn these observations into an algorithm that takes post-trained models and a new reward function, and approximates the target RL policy without actually running any additional RL training. We experimentally validate our approach across synthetic and real rewards, spanning both text and image modalities.