发表机构
Massachusetts Institute of Technology; MIT-IBM Computing Research Lab; Stanford University; University of California, San Diego(麻省理工学院; 麻省理工学院-IBM计算研究实验室; 斯坦福大学; 加利福尼亚大学圣地亚哥分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对强化学习用于视觉-语言-动作模型后期训练时样本效率低的问题,提出事后诸葛亮式学习方法,通过对失败轨迹重标记,让策略联合原始与重标记轨迹训练,在分布外任务上显著提升样本效率并优于基线。
AI 中文摘要
强化学习(RL)越来越多地用于视觉-语言-动作(VLA)模型的后期训练,但每次更新都需要消耗机器人的轨迹,而这些轨迹收集起来缓慢且成本高昂,因此样本效率成为核心问题。操纵任务通常只提供稀疏奖励,导致弱策略在训练早期几乎每次轨迹都失败且几乎无学习价值。我们提出了事后诸葛亮式学习(LfH),通过根据失败轨迹实际达成的任务对其进行评分,将事后重标记引入VLA的RL后期训练。单个视觉-语言模型对指令和奖励进行重标记,为一组失败轨迹提出事后诸葛亮式指令并评分,策略在重标记和原始轨迹上联合训练。在分布外的LIBERO-PRO任务上,标准RL进展缓慢,而LfH实现了样本效率5倍的提升,并优于密集进展奖励基线。这些增益在不同VLA主干和实体Franka机器人上均成立。
英文摘要
Reinforcement learning is increasingly used to fine-tune vision-language-action (VLA) models, but robot interaction is expensive and learning becomes highly sample inefficient when successful rollouts are rare. When reward is assigned only for completing the commanded task, a failed rollout is treated as having no value even if it successfully executes behaviors relevant to that task. A robot that fails to place the correct object in a bowl may still move that object toward the bowl or place a different object inside it, demonstrating objects and actions that can be reused to solve the target task. These behaviors define auxiliary tasks that the policy can already solve, providing useful learning signals even before it can solve the harder target task. We introduce $\textit{Learning from Hindsight (LfH)}$, which turns such failures into additional learning signals. Using a pretrained vision-language model, LfH relabels failed rollouts with the behaviors they actually accomplish and trains the policy jointly on the commanded task and these auxiliary tasks. On out-of-distribution LIBERO-PRO manipulation tasks, LfH matches the final performance of GRPO with approximately $5\times$ fewer rollouts and improves sample efficiency across multiple VLA backbones. On a physical Franka robot, LfH raises success from $0\%$ to $56\%$ within 160 training rollouts, while GRPO reaches $22\%$.