arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非所有轨迹都值得学习:关于后训练强化学习中的轨迹估值

Not All Rollouts Are Worth Learning: On Trajectory Valuation for Post-Training Reinforcement Learning

Xuesong Jia, Ziao Yang, Zhanhe Huang, Hongfu Liu

arXiv 2609.35072首次发表:更新:

发表机构

Brandeis University(布兰迪斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对强化学习在线训练中轨迹质量参差不齐的问题,提出动态轨迹估值(DTV)框架,仅利用梯度信息在小批量级别评估并过滤有害轨迹,以最小开销集成至现有流程,在PPO、GRPO和DPO等设置中提升性能、数据效率与优化稳定性。

AI 中文摘要

我们考虑强化学习中的轨迹估值问题:如何在在线训练过程中识别并减轻有害轨迹的影响。与分类任务中数据估值依赖于固定的训练集和验证集不同,强化学习涉及动态生成的轨迹,且缺乏明确的验证信号,这使得传统的基于影响力的方法无法适用。我们提出了动态轨迹估值(DTV),一个简单而高效的框架,它在小批量级别估计轨迹效用,并仅基于梯度信息过滤有害轨迹。通过在优化层面操作,DTV能够以最小的开销无缝集成到现有的强化学习流程中。在包括PPO、GRPO和DPO在内的多种设置下进行的广泛实验表明,DTV持续提升性能、增强数据效率并稳定优化过程。

英文摘要

We consider the problem of trajectory valuation in reinforcement learning: how to identify and mitigate detrimental trajectories during online training. Unlike classification, where data valuation relies on fixed training and validation sets, reinforcement learning involves dynamically generated trajectories without explicit validation signals, making conventional influence-based methods inapplicable. We propose Dynamic Trajectory Valuation (DTV), a simple and efficient framework that estimates trajectory utility at the mini-batch level and filters detrimental trajectories based solely on gradient information. By operating at the optimization level, DTV integrates seamlessly with existing reinforcement learning pipelines with minimal overhead. Extensive experiments across diverse settings, including PPO, GRPO, and DPO, demonstrate that DTV consistently improves performance, enhances data efficiency, and stabilizes optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑