arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27068cs.RO

HiRE:用于策略微调的事后奖励编辑

HiRE: Hindsight Reward Editing for Policy Finetuning

Haoyi Niu, Zhengtao Han, Yufeng Ji, Zhongyu Li, Koushil Sreenath

首次发表
浏览论文内容

中文总结 AI 辅助

HiRE通过事后对比成功与失败轨迹,校准基础表征模型并惩罚陷阱状态,提供密集控制感知奖励,提升机器人策略微调的样本效率和性能。

中文摘要 AI 辅助

预训练的机器人策略总是需要微调以适应特定环境。强化学习(RL)提供了高性能潜力,因为它改善动作的最优性,而非简单地模仿数据。然而,这种潜力在很大程度上依赖于奖励质量。稀疏奖励缺乏过程反馈,人工设计的奖励成本高昂且有偏差,而来自基础表征的语义奖励通常不是以控制为中心的。我们提出事后奖励编辑(HiRE),一种免训练框架,以打破这一奖励瓶颈。HiRE通过事后对比成功与失败轨迹,将基础表征模型的广泛知识与物理控制意识相连接。它通过识别被预测为高奖励状态但最终导致失败的“陷阱状态”,反之亦然,来校准基础表征模型。HiRE明确惩罚这些陷阱,同时提升关键成功状态的奖励。该方法可灵活兼容任何基础表征和RL算法。实验表明,HiRE通过提供密集、控制感知的反馈,防止价值函数崩溃和奖励黑客,持续优于其他奖励方案,从而实现卓越的样本效率、稳定的策略更新和更高的性能上限,例如,至少为基础策略性能的3倍。定性结果见https://hire-project.github.io。

英文摘要

Pre-trained robot policies always require finetuning to adapt to specific environments. Reinforcement Learning (RL) offers high performance potential because it improves action optimality rather than simply mimicking data. However, such potential depends heavily on reward quality. Sparse rewards lack process feedback, human-designed rewards are costly and biased, and semantic rewards from foundation representations are often not control-centric. We propose Hindsight Reward Editing (HiRE), a training-free framework to break this reward bottleneck. HiRE bridges the broad knowledge of foundation representation models with physical control awareness, by contrasting successful and failed trajectories in hindsight. It calibrates foundation representation models by identifying "trap states" that are predicted as high-rewarding states yet eventually result in failure, and vice versa. HiRE explicitly penalizes these traps while boosting rewards for critical successful states. This approach can be flexibly compatible with any foundation representations and RL algorithms. Experiments show that HiRE consistently outperforms other reward recipes by delivering dense, control-aware feedback that prevents value function collapse and reward hacking, thereby achieving superior sample efficiency, stable policy updates, and higher performance ceilings, e.g., at least 3x performance of the base policies. Qualitative results are at https://hire-project.github.io .

发表机构

  • University of California, Berkeley(加州大学伯克利分校)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑