arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13026cs.RO

时间差分GRPO:超越视觉-语言-动作强化学习中的轨迹级信用分配

Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning

Yao Zhou, Hang Gao, Fengge Wu, Changwen Zheng, Wenwen Qiang

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉-语言-动作强化学习中轨迹级信用混叠问题,提出时间差分GRPO方法,在RoboTwin 2.0和LIBERO-Long上提升了任务性能与样本效率。

中文摘要 AI 辅助

结果驱动的强化学习提供了一种可扩展的方式,可从稀疏的任务成功反馈中对视觉-语言-动作(VLA)策略进行后训练。在常见的基于GRPO的VLA后训练中,一次回合的优势会应用于轨迹中的每个动作。一个完成了多个有效阶段但在后期失败的回合,可能会惩罚产生早期进展的动作,我们将此称为轨迹级信用混叠。时间差分GRPO通过构建可检测的任务阶段、将每个回合与阶段特定的动作区间对齐、仅比较进入相同阶段的回合来解决该问题。所得的阶段优势会应用于单次策略更新中对应的区间。在RoboTwin 2.0上,时间差分GRPO提高了任务成功率和样本效率,在不同任务时长下均有稳定提升;在LIBERO-Long上进行的受控更新保留了共享的前提阶段,并将改进集中在回合结果出现分歧的第一阶段。

英文摘要

Outcome-driven reinforcement learning offers a scalable way to post-train vision-language-action (VLA) policies from sparse task-success feedback. In common GRPO-based VLA post-training, one rollout-level advantage is applied to every action in the trajectory. A rollout that completes several valid stages but fails later can therefore penalize the actions that produced its earlier progress. We call this trajectory-level credit aliasing. Temporal GRPO addresses this problem by constructing detectable task stages, aligning each rollout with stage-specific action intervals, and comparing only rollouts that have entered the same stage. The resulting stage advantages are applied to their corresponding intervals in a single policy update. On RoboTwin 2.0, Temporal GRPO improves task success and sample efficiency, with consistent gains across task horizons. Controlled updates on LIBERO-Long preserve shared prerequisite stages and concentrate improvement at the first stage where rollout outcomes diverge.

↑