arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BCPPO:受巴舍利耶启发的用于尾部风险感知安全强化学习的约束近端策略优化算法

BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning

Dongsheng Hou, Yanqiao Chen, Yuhan Rui

arXiv 2608.30283首次发表:更新:

发表机构

Southern University of Science and Technology(南方科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出BCPPO算法,通过带随机样本掩码的成本预测网络分歧结合巴舍利耶公式生成平滑惩罚,在175次运行的多任务测试中,其平均回报更高、平均CVaR更低,实现了奖励与安全的实用平衡。

AI 中文摘要

期望成本约束仍可能允许罕见的高成本事件发生。蒙特卡洛条件风险价值(CVaR)梯度在高置信度下可能存在噪声,而对结果分布进行建模的评论者网络会增加复杂度。本文提出BCPPO(Bachelier-Inspired Constrained Proximal Policy Optimization,受巴舍利耶启发的约束近端策略优化算法),这是一种基于近端策略优化(PPO)的方法。通过随机样本掩码训练、单独初始化的成本预测网络(评论者)会产生分歧,这种分歧可标记出对训练数据中出现的状态-动作区域以及评论者训练敏感的预测。用于计算参考水平以上期望数量的巴舍利耶公式,将这种分歧转化为平滑的策略更新惩罚。该惩罚的梯度不会改变评论者,因此时间差分(TD)评论者学习保持不变。饱和感知控制器会调整平均成本惩罚,并在该惩罚被裁剪时阻止累积误差增长。部署阶段仅保留策略网络。分歧惩罚既不是尾部事件概率,也不是保证误差边界,无法提供安全保证。在具有共享任务、成本、预算、训练步数和评估种子的175次运行中,没有任何对比算法在任何任务中同时达到比BCPPO更高的平均回报和更低的平均CVaR。在Push1任务上,BCPPO的回报不低于任何对比算法,CVaR不高于任何对比算法,且至少存在一项严格优势。这些结果支持在奖励、对经训练的评论者预测变化的谨慎态度以及仅策略部署之间实现实用平衡。

英文摘要

Expected-cost constraints can still permit rare, high-cost events. Monte Carlo conditional value at risk (CVaR) gradients can be noisy at high confidence, whereas critics that model an outcome distribution add complexity. We propose BCPPO (Bachelier-Inspired Constrained Proximal Policy Optimization), a proximal policy optimization (PPO) method. Separately initialized cost-prediction networks (critics), trained with random sample masks, produce disagreement that marks predictions sensitive to which state-action regions occur in the training data and to critic training. A Bachelier formula for the expected amount above a reference level converts this disagreement into a smooth policy-update penalty. Gradients from this penalty do not alter the critics, so temporal-difference (TD) critic learning is unchanged. A saturation-aware controller adjusts the mean-cost penalty and stops accumulated error from growing while that penalty is clipped. Deployment retains only the policy network. The disagreement penalty is neither a tail-event probability nor a guaranteed error bound, and it provides no safety guarantee. Across 175 runs with shared tasks, costs, budgets, training steps, and evaluation seeds, no comparator attains both higher mean return and lower mean CVaR than BCPPO in any task. On Push1, BCPPO has no lower return and no higher CVaR than every comparator, with at least one strict gain. These results support a practical balance among reward, caution around cost predictions that vary across trained critics, and policy-only deployment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑