arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26571cs.LGcs.RO

到达并生存:基于一位失败信号扩展安全目标条件策略学习

Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals

  • Southeast University(东南大学)
  • Yinwang Intel. Tech. Co. Ltd.(银网智能科技有限公司)
  • Lab(2030实验室)

机构由 AI 辅助整理,请以论文原文为准。

Guopeng Li, Yiyang Duan, Yiru Jiao, Chengcheng Xu

AI总结:

本文针对失败终止CRL的系统性偏差,提出Safe-CRL方法,仅用失败的一位信号,在12个机器人任务中提升了存活率与目标到达性能,完善了失败终止下的CRL理论。

AI中文摘要:

对比强化学习(CRL)通过将策略学习转化为自监督对比目标,可在目标条件任务中有效扩展。然而在失败终止马尔可夫决策过程中,现有CRL在构建正样本时仅考虑失败前的未来目标,未考虑失败终止移除的概率质量。理论分析表明,该遗漏会导致目标到达值出现系统性高估偏差,进而使接近失败的轨迹提供不成比例的强成功监督,尽管其未来占据空间很小。不安全动作会通过灾难性失败自举得到强化,导致策略学习失败和不可持续的目标到达行为。为解决此问题,本文提出两个最小但有效的修正:质量加权InfoNCE修正评论者学习中对短存活未来的过度加权,对数存活质量分数在策略优化中恢复缺失的存活质量。由此得到的方法为安全对比强化学习(Safe-CRL),仅需失败终止提供的一位信号即可扩展安全目标条件策略学习。在12个易失败的机器人导航和运动任务中,Safe-CRL始终提高存活率,在目标到达性能上显著优于Scaling-CRL基线,且深度Safe-CRL策略表现出复杂的避失败行为。本研究完善了失败终止下的CRL理论,提供了可扩展的安全RL框架,代码可通过该https URL获取。

英文摘要:

Contrastive reinforcement learning (CRL) scales effectively in goal-conditioned tasks by casting policy learning into a self-supervised contrastive objective. However, in a failure-terminated Markov decision process, established CRL considers pre-failure future goals only when constructing positive samples, without accounting for the probability mass removed by failure termination. Our theoretical analysis shows that this omission induces a systematic overestimation bias in goal-reaching values. Consequently, near-failure trajectories provide disproportionately strong supervision of success despite retaining little future occupancy. Unsafe actions can thereby be reinforced through catastrophic failure bootstrapping, leading to failed policy learning and unsustainable goal-reaching behaviours. To address this problem, we introduce two minimal yet strong corrections: mass-weighted InfoNCE corrects the overweighting of short surviving futures in critic learning, and a log-survival-mass score restores the missing survival mass in policy optimization. The resulting method, Safe Contrastive Reinforcement Learning (Safe-CRL), requires only the one-bit signal provided by failure termination to scale safe goal-conditioned policy learning. Across twelve failure-prone robot navigation and locomotion tasks, Safe-CRL consistently improves survival and substantially outperforms the Scaling-CRL baseline in goal-reaching performance. Additionally, deep Safe-CRL policies exhibit complex failure-avoidance behaviours. This study completes the CRL theory under failure termination and provides a scalable safe RL framework. The code is available via https://github.com/RomainLITUD/safe-crl.

补充信息

↑