尾似然强化学习
Tail-Likelihood Reinforcement Learning
- Carnegie Mellon University(卡内基梅隆大学)
- University of California, Berkeley(加州大学伯克利分校)
- Impossible, Inc.(Impossible公司)
- Together AI
- Aurora Innovation
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对强化学习优化平均奖励的局限,提出TailRL方法,通过最大化超过随机奖励阈值的对数概率,利用稀有高奖励样本提升模型性能。
AI中文摘要:
强化学习通常优化平均奖励。对于生成式策略,平均值会掩盖一个重要差异:两个策略可以获得相同的平均奖励,但产生稀有高奖励轨迹的概率却大不相同。这一点在训练和推理过程中采样量增加时至关重要,因为采样的好处取决于保留高奖励结果的概率质量。我们提议直接优化这种覆盖范围。我们不仅考虑预期奖励,还考虑其所有上尾部分:对于每个奖励阈值,策略超过该阈值的概率是多少?这将连续奖励转化为一系列二元成功事件。我们引入尾似然强化学习(TailRL),它最大化超过随机选择的奖励阈值的对数概率。其梯度对稀有高奖励轨迹赋予更多权重,可解释为Best-of-(k)梯度的混合。TailRL仅需对优势函数进行简单修改,即可与现有强化学习流程兼容。在目标定位、迷宫导航、GUI定位和代码优化任务中,TailRL利用稀有高奖励训练样本避免次优解,生成的模型在推理时能从更多样本中获得更多收益。
英文摘要:
Reinforcement learning typically optimizes average reward. For generative policies, the average can hide an important distinction: two policies can achieve the same mean reward while having very different chances of producing a rare but high-reward rollout. This matters as sampling increases during training and inference, since its benefit depends on retaining probability mass on high-reward outcomes. We propose to optimize this coverage directly. Rather than considering only expected reward, we consider all of its upper tails: for each reward threshold, how likely is the policy to exceed it? This turns a continuous reward into a family of binary success events. We introduce Tail-Likelihood Reinforcement Learning (TailRL), which maximizes the log-probability of exceeding a randomly chosen reward threshold. Its gradient gives more weight to rare, high-reward rollouts and can be interpreted as a mixture of Best-of-k gradients. TailRL requires only a simple modification to the advantage function, making it compatible with existing reinforcement learning pipelines. Across object localization, maze navigation, GUI grounding, and code optimization, TailRL leverages rare high-reward training samples to avoid suboptimal solutions and yields models that benefit more from additional samples at inference time.