arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38890cs.RO

PRICE 动作块:具身强化学习的物理关系信用分配

PRICE the Action Chunks: Physical Relational Credit Assignment for Embodied Reinforcement Learning

Yangang Zou, Jiajun Lu, Weitao Zhou, Haibao Yu, Bozhou Zhang, Jiawei Wang, Honglong Tian, Minglei Li, Li Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

提出PRICE方法,利用跨轨迹的物理关系图从终端结果中推断动作块级信用,无需辅助评估器,在具身强化学习中提升任务成功率并加速学习。

中文摘要 AI 辅助

基于结果(Outcome-based)的强化学习(RL)使用终端成功信号对视觉-语言-动作策略进行后训练,但将相同的轨迹级优势(trajectory-level advantage)分配给每个动作块(action chunk)。因此,一个失败的回合可能会惩罚有用的早期动作,仿佛它们导致了失败。现有方法通过学习的评估器寻求更细粒度的反馈,这增加了任务特定的监督或额外的模型训练。据我们所知,我们首次探索了跨轨迹的物理关系是否可以在仅依赖终端结果、无需辅助评估器的情况下,为具身强化学习中的动作块提供信用分配。关键洞察在于,达到相应物理状态的轨迹(rollouts)可以相互作为参考:它们的终端结果为评估局部进展提供了证据。我们提出了从回合中推断信用的物理关系(PRICE),包含两个组件:(i)一个物理关系图,在相应的块边界汇集当前和历史结果以估计成功潜力;(ii)置信门控的信用分配,利用这些潜力的变化来细化轨迹级监督。我们的分析将预言机潜力变化与终端成功目标联系起来,并为结果无关的证据池提供了有限样本方向性界限。独立的延续测试表明,PRICE保留的信用与局部进展一致,而在LIBERO、RoboTwin 2.0和真实机器人上的实验表明,与基于结果的基线相比,任务成功率有所提高且学习速度更快。

英文摘要

Outcome-based reinforcement learning (RL) post-trains vision--language--action policies using terminal success signals, but assigns the same trajectory-level advantage to every action chunk. A failed episode can thus penalize useful early actions as if they caused the failure. Existing approaches seek finer-grained feedback through learned evaluators, adding task-specific supervision or additional model training. We explore, for the first time to our knowledge, whether physical relations across trajectories can provide action-chunk credit in embodied RL from terminal outcomes alone, without an auxiliary evaluator. The key insight is that rollouts reaching corresponding physical situations can serve as references for one another: their terminal outcomes provide evidence for assessing local progress. We introduce Physical Relations for Inferring Credit from Episodes(PRICE), with two components: (i) a physical relational graph that pools current and historical outcomes at corresponding chunk boundaries to estimate success potentials; and (ii) confidence-gated credit assignment that uses changes in these potentials to refine trajectory-level supervision. Our analysis connects oracle potential changes to the terminal-success objective and provides a finite-sample directional bound for outcome-independent evidence pools. Independent continuation tests show that PRICE's retained credits align with local progress, while experiments on LIBERO, RoboTwin 2.0, and real robots demonstrate improved task success over outcome-based baselines and faster learning.

发表机构

  • Fudan University(复旦大学)
  • Shanghai Innovation Institute(上海创新研究院)
  • Beihang University(北京航空航天大学)
  • Simple AI
  • Tuojing Intelligence(拓境智能)
  • Tsinghua University(清华大学)
  • The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

↑