arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02987cs.LGstat.ML

尾似然强化学习

Tail-Likelihood Reinforcement Learning

  • Carnegie Mellon University(卡内基梅隆大学)
  • University of California, Berkeley(加州大学伯克利分校)
  • Impossible, Inc.(Impossible公司)
  • Together AI
  • Aurora Innovation

机构由 AI 辅助整理,请以论文原文为准。

Shrinivas Ramasubramanian, Daman Arora, Fahim Tajwar, Guanning Zeng, Qingyang Wu, Zhongzhu Zhou, Chenfeng Xu, Haiwen Feng, Yuda Song, Aarti Singh, Ruslan Salakh… 展开作者

Shrinivas Ramasubramanian, Daman Arora, Fahim Tajwar, Guanning Zeng, Qingyang Wu, Zhongzhu Zhou, Chenfeng Xu, Haiwen Feng, Yuda Song, Aarti Singh, Ruslan Salakhutdinov, J. Andrew Bagnell, Jeff Schneider, Andrea Zanette

AI总结:

针对强化学习优化平均奖励的局限,提出TailRL方法,通过最大化超过随机奖励阈值的对数概率,利用稀有高奖励样本提升模型性能。

AI中文摘要:

强化学习通常优化平均奖励。对于生成式策略,平均值会掩盖一个重要差异:两个策略可以获得相同的平均奖励,但产生稀有高奖励轨迹的概率却大不相同。这一点在训练和推理过程中采样量增加时至关重要,因为采样的好处取决于保留高奖励结果的概率质量。我们提议直接优化这种覆盖范围。我们不仅考虑预期奖励,还考虑其所有上尾部分:对于每个奖励阈值,策略超过该阈值的概率是多少?这将连续奖励转化为一系列二元成功事件。我们引入尾似然强化学习(TailRL),它最大化超过随机选择的奖励阈值的对数概率。其梯度对稀有高奖励轨迹赋予更多权重,可解释为Best-of-(k)梯度的混合。TailRL仅需对优势函数进行简单修改,即可与现有强化学习流程兼容。在目标定位、迷宫导航、GUI定位和代码优化任务中,TailRL利用稀有高奖励训练样本避免次优解,生成的模型在推理时能从更多样本中获得更多收益。

英文摘要:

Reinforcement learning typically optimizes average reward. For generative policies, the average can hide an important distinction: two policies can achieve the same mean reward while having very different chances of producing a rare but high-reward rollout. This matters as sampling increases during training and inference, since its benefit depends on retaining probability mass on high-reward outcomes. We propose to optimize this coverage directly. Rather than considering only expected reward, we consider all of its upper tails: for each reward threshold, how likely is the policy to exceed it? This turns a continuous reward into a family of binary success events. We introduce Tail-Likelihood Reinforcement Learning (TailRL), which maximizes the log-probability of exceeding a randomly chosen reward threshold. Its gradient gives more weight to rare, high-reward rollouts and can be interpreted as a mixture of Best-of-k gradients. TailRL requires only a simple modification to the advantage function, making it compatible with existing reinforcement learning pipelines. Across object localization, maze navigation, GUI grounding, and code optimization, TailRL leverages rare high-reward training samples to avoid suboptimal solutions and yields models that benefit more from additional samples at inference time.

↑