安全强化学习的评估指标
Evaluation Metrics for Safe Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对现有安全强化学习评估仅关注平均安全性的不足,提出涵盖违反频率、严重性及跨任务一致性的新评估指标和安全等级系统,并基于多导航任务实证验证,推荐联合报告聚合、分布及具体结果,同时开源SafeRLEval套件。
中文摘要 AI 辅助
安全强化学习(RL)通常被形式化为约束马尔可夫决策过程(CMDP),其中智能体在最大化期望奖励的同时,需将其期望累积成本保持在指定的安全界限以下。现有的安全强化学习基准主要遵循这种基于期望的保证,报告算法在平均意义上是否安全。我们认为,这一惯例不足以可靠地表征算法的真实安全性:它未能捕捉安全界限被违反的频率和严重程度,未能反映这种安全性是否在不同任务和安全界限间保持一致,也未能反映训练期间的行为是否能代表最终收敛策略的行为。因此,我们引入了(i)针对安全强化学习的评估指标,以解决上述每一个问题,并允许跨任务和安全界限进行聚合。此外,我们定义了(ii)一个安全等级系统,以系统地对算法在训练阶段和最终策略上的安全性和可靠性进行分类和比较。利用这一框架,我们提供了(iii)跨多个安全导航任务的实证安全评估。我们的结果表明,聚合指标、分布报告以及针对任务和安全界限的具体结果,各自揭示了其他指标无法提供的信息。因此,我们建议将这三者联合报告,而非像常见做法那样将这些信息压缩为单一数值。我们提供了SafeRLEval,一个开源评估套件,以支持未来安全强化学习研究中可靠的安全表征。
英文摘要
Safe reinforcement learning (RL) is commonly formalized as a Constrained Markov Decision Process (CMDP), in which an agent maximizes expected reward while keeping its expected cumulative cost below a specified safety bound. Existing safe RL benchmarks predominantly report whether an algorithm is safe on average, following this expectation-based guarantee. We argue that this convention is insufficient to reliably characterize an algorithm's true safety: it fails to capture how often and how severely the safety bound is violated, whether this holds consistently across tasks and safety bounds, and whether training-time behavior is representative of behavior of the final converged policy. Therefore, we introduce (i) evaluation metrics for safe RL that address each of these concerns and in addition allow for aggregation across tasks and safety bounds. We furthermore define (ii) a safety tier system to systematically categorize and compare algorithms in terms of safety and reliability at both training and for a final policy. Using this framework, we provide (iii) an empirical safety evaluation across multiple safety navigation tasks. Our results show that aggregate metrics, distributional reporting, and task- and safety bound-specific results each reveal information the other metrics cannot. We therefore recommend reporting all three jointly, rather than compressing this information into a single value, as is common practice. We provide SafeRLEval, an open-source evaluation suite to support the reliable characterization of safety in future safe RL research.
发表机构
- Leiden University(莱顿大学)
机构由 AI 辅助整理,请以论文原文为准。