PRICE 动作块:具身强化学习的物理关系信用分配
PRICE the Action Chunks: Physical Relational Credit Assignment for Embodied Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
提出PRICE方法,利用跨轨迹的物理关系图从终端结果中推断动作块级信用,无需辅助评估器,在具身强化学习中提升任务成功率并加速学习。
中文摘要 AI 辅助
基于结果(Outcome-based)的强化学习(RL)使用终端成功信号对视觉-语言-动作策略进行后训练,但将相同的轨迹级优势(trajectory-level advantage)分配给每个动作块(action chunk)。因此,一个失败的回合可能会惩罚有用的早期动作,仿佛它们导致了失败。现有方法通过学习的评估器寻求更细粒度的反馈,这增加了任务特定的监督或额外的模型训练。据我们所知,我们首次探索了跨轨迹的物理关系是否可以在仅依赖终端结果、无需辅助评估器的情况下,为具身强化学习中的动作块提供信用分配。关键洞察在于,达到相应物理状态的轨迹(rollouts)可以相互作为参考:它们的终端结果为评估局部进展提供了证据。我们提出了从回合中推断信用的物理关系(PRICE),包含两个组件:(i)一个物理关系图,在相应的块边界汇集当前和历史结果以估计成功潜力;(ii)置信门控的信用分配,利用这些潜力的变化来细化轨迹级监督。我们的分析将预言机潜力变化与终端成功目标联系起来,并为结果无关的证据池提供了有限样本方向性界限。独立的延续测试表明,PRICE保留的信用与局部进展一致,而在LIBERO、RoboTwin 2.0和真实机器人上的实验表明,与基于结果的基线相比,任务成功率有所提高且学习速度更快。
英文摘要
Outcome-based reinforcement learning (RL) post-trains vision--language--action policies using terminal success signals, but assigns the same trajectory-level advantage to every action chunk. A failed episode can thus penalize useful early actions as if they caused the failure. Existing approaches seek finer-grained feedback through learned evaluators, adding task-specific supervision or additional model training. We explore, for the first time to our knowledge, whether physical relations across trajectories can provide action-chunk credit in embodied RL from terminal outcomes alone, without an auxiliary evaluator. The key insight is that rollouts reaching corresponding physical situations can serve as references for one another: their terminal outcomes provide evidence for assessing local progress. We introduce Physical Relations for Inferring Credit from Episodes(PRICE), with two components: (i) a physical relational graph that pools current and historical outcomes at corresponding chunk boundaries to estimate success potentials; and (ii) confidence-gated credit assignment that uses changes in these potentials to refine trajectory-level supervision. Our analysis connects oracle potential changes to the terminal-success objective and provides a finite-sample directional bound for outcome-independent evidence pools. Independent continuation tests show that PRICE's retained credits align with local progress, while experiments on LIBERO, RoboTwin 2.0, and real robots demonstrate improved task success over outcome-based baselines and faster learning.
发表机构
- Fudan University(复旦大学)
- Shanghai Innovation Institute(上海创新研究院)
- Beihang University(北京航空航天大学)
- Simple AI
- Tuojing Intelligence(拓境智能)
- Tsinghua University(清华大学)
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。