Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals
到达并生存:基于一位失败信号扩展安全目标条件策略学习
机构 * Southeast University(东南大学) ; Yinwang Intel. Tech. Co. Ltd.(银网智能科技有限公司) ; Lab(2030实验室)
专题命中 BEV与占用 :occupancy(abstract);分类 cs.RO
AI总结 本文针对失败终止CRL的系统性偏差,提出Safe-CRL方法,仅用失败的一位信号,在12个机器人任务中提升了存活率与目标到达性能,完善了失败终止下的CRL理论。
Comments 21 pages, 14 figures, 5 tables, Code: this https URL (https://github.com/RomainLITUD/safe-crl)