发表机构
McGill University; Mila; Technical University of Munich; Université de Montréal(麦吉尔大学; 米拉; 慕尼黑工业大学; 蒙特利尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在物理机器人上训练强化学习智能体减少跌倒问题,提出对PPO改进的SafeExplorer算法,核心是无偏策略梯度估计器,还添加加速学习组件,实验表明该算法大幅减少训练时间跌倒且奖励表现良好。
AI 中文摘要
在物理机器人上直接训练强化学习智能体,每次跌倒成本高昂,因此目标是在训练期间尽量减少跌倒,而非像约束马尔可夫决策过程公式那样权衡跌倒与回报。标准缓解措施是当智能体离开指定安全区域时将控制权交给单独的恢复策略,但这会导致混合策略推出对策略更新产生偏差,且恢复策略为确定性时重要性采样校正不明确。我们通过对近端策略优化(PPO)进行改进来解决此偏差,其核心是无偏策略梯度估计器,仅在安全时间步使用得分函数且不评估恢复策略密度。此外,因恢复策略在安全区域边界附近信用分配慢,还添加两个组件加速学习。在三个环境、五组种子的基准测试中,该算法在HalfCheetah、Ant和Unitree Go1上比标准PPO将训练时间跌倒减少233倍、48倍和26倍,同时匹配或超过PPO的最终奖励,在Ant上,恢复策略不可靠时,它是唯一达到最佳最终奖励80%的方法。
英文摘要
Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do. A standard mitigation hands control to a separate recovery policy whenever the agent leaves a designer-specified safe region (a subset of state space it should stay within), but the resulting mixed-policy rollouts silently bias every on-policy update, and the importance-sampling correction that would remove this bias is ill-defined whenever the recovery policy is deterministic. We address this bias with a drop-in modification of proximal policy optimization (PPO). Its core is an unbiased policy-gradient estimator that uses the score function only at safe timesteps and never evaluates the recovery policy's density, so it stays valid even when the recovery policy is deterministic, exactly where importance sampling breaks, and it empirically dominates importance sampling even when the recovery policy is stochastic. Because the recovery policy still makes credit assignment slow near the safe-region boundary, two further components accelerate learning: a closed-form value for recovery-triggering states when dynamics and recovery are deterministic, and an imitation loss that copies recovery actions only when recovery succeeds. On a three-environment, five-seed benchmark, the resulting algorithm reduces training-time falls by factors of 233x, 48x, and 26x on HalfCheetah, Ant, and Unitree Go1 over standard PPO, while matching or exceeding PPO's final reward, and on Ant, where the recovery policy is unreliable, it is the only method that reaches 80% of the best final reward.