Prism-GRPO:通过拆分相同结果组实现更快的视觉-语言-动作(VLA)策略优化
Prism-GRPO: Faster VLA Policy Optimization via Splitting Same-outcome Groups
浏览论文内容
中文总结 AI 辅助
该研究提出 Prism-GRPO 算法,通过拆分 GRPO 的相同结果组并引入加权执行质量分数,减少机器人 rollout 预算浪费,在四个 RoboTwin 任务中使达到目标成功率的 rollout 最多减少 56%,还抑制了奖励黑客捷径。
中文摘要 AI 辅助
GRPO因无需训练评论家(critic),正越来越多地用于视觉-语言-动作(VLA)策略的强化学习,这一简化带来了采样成本:组相对优势需要从每个场景进行多次 rollout。在二元成功奖励下,所有 rollout 均成功或均失败的组具有零优势,会被动态采样丢弃。这些组在训练初期尤其常见,此时大多数 rollout 失败,浪费了大量昂贵的机器人 rollout 预算。我们提出 Prism-GRPO,它通过加权轨迹级执行质量分数增强二元结果奖励。通过将相同结果组拆分为质量谱,Prism-GRPO 在确保每个成功仍优于每个失败的同时,恢复了训练信号。质量分数可从模拟器接触、执行的动作或视觉观测中推导,无需特定任务的进度奖励。我们证明 Prism-GRPO 绝不会增加因具有零优势而丢弃采样组的概率,并推导了梯度对齐条件,在该条件下其组合更新仍是任务成功的局部上升方向。在四个涵盖不同时间范围和协调模式的 RoboTwin 任务中,Prism-GRPO 在匹配的 rollout 预算下提高了成功率和质量,且达到目标成功率所需的 rollout 最多减少了 56%。它还抑制了奖励黑客捷径,更干净的行为可直接部署到真实机器人上。通过 ablation 实验,我们展示了其在基于接触、平滑度和 VLM 的质量信号上的一致增益。
英文摘要
GRPO is increasingly used for reinforcement learning of vision-language-action (VLA) policies because, unlike PPO, it does not require training a critic. This simplification comes with a sampling cost: group-relative advantages require multiple rollouts from each scene. Under binary success rewards, groups whose rollouts all succeed or all fail have zero advantage and are discarded by dynamic sampling. These groups are especially common early in training, when most rollouts fail, wasting much of the expensive robotic rollout budget. We introduce Prism-GRPO, which augments binary outcome reward with a weighted trajectory-level execution-quality score. By splitting same-outcome groups into a quality spectrum, Prism-GRPO recovers training signal while ensuring that every success still outranks every failure. Quality scores can be derived from simulator contacts, executed actions, or visual observations, avoiding task-specific progress rewards. We prove that Prism-GRPO never increases the probability that a sampled group is discarded for having zero advantages, and derive a gradient-alignment condition under which its combined update remains a local ascent direction for task success. Across four RoboTwin tasks spanning different horizons and coordination patterns, Prism-GRPO improves success and quality at matched rollout budgets and reaches target success rates with up to 56% fewer rollouts. It also suppresses a reward-hacking shortcut, with the cleaner behavior transferring under direct deployment to a real robot. Through ablations, we show consistent gains across contact-, smoothness-, and VLM-derived quality signals.
发表机构
- AWS AI(亚马逊云科技人工智能部门)
机构由 AI 辅助整理,请以论文原文为准。