鲁棒峰值成本约束强化学习
Robust Peak-cost Constrained Reinforcement Learning
查看机构详情
- Washington State University(华盛顿州立大学)
- The Ohio State University(俄亥俄州立大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究鲁棒峰值成本约束强化学习,针对安全关键应用。开发代理优化框架和鲁棒值估计方法,解决模拟器到现实世界的不匹配,在动态扰动下加强安全性并保持奖励性能。
中文摘要 AI 辅助
我们研究鲁棒峰值成本约束强化学习(RP-CRL),目标是在控制轨迹中遇到的最大成本的同时最大化预期奖励。这种设置受安全关键应用的推动,在这些应用中单个大的违规可能是灾难性的,标准基于预期累积成本的CMDP框架无法充分捕捉。现有可达性约束强化学习方法采用基于拉格朗日的方法,但峰值成本约束MDP的潜在对偶性质仍不清楚。我们表明,与标准CMDP不同,峰值成本约束MDP可能不存在零对偶间隙。我们进一步考虑一种鲁棒公式来解决过渡动态中的模拟器到现实世界的不匹配。为解决此问题,我们开发了一个代理优化框架和基于积分概率度量的鲁棒值估计方法。我们证明,通过适当选择超参数,代理解决方案获得与原始问题相同的鲁棒奖励值,同时违反约束最多为ε。实验表明,该方法在动态扰动下有效加强了安全性,同时保持了强大的奖励性能。
英文摘要
We study robust peak-cost constrained reinforcement learning (\ours), where the objective is to maximize expected reward while controlling the maximum cost encountered along a trajectory. This setting is motivated by safety-critical applications in which a single large violation can be catastrophic and therefore cannot be adequately captured by the standard CMDP framework based on expected cumulative cost. Existing reachability-constrained RL methods adopt Lagrangian-based approaches, yet the underlying duality properties of peak-cost constrained MDPs remain unclear. We show that, unlike standard CMDPs, peak-cost constrained MDPs may not admit zero duality gap. We further consider a robust formulation to address simulator-to-real-world mismatch in the transition dynamics. To solve this problem, we develop a surrogate optimization framework and a robust value estimation method based on integral probability metrics. We prove that, with appropriate hyperparameter choices, the surrogate solution attains the same robust reward value as the original problem while violating the constraint by at most \(ε\). Experiments show that the proposed method effectively enforces safety under dynamics perturbations while retaining strong reward performance across diverse environments.