发表机构
McMaster University(麦克马斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究非平稳强化学习中的安全问题,提出以调整速度为安全约束,用上下文表示和预测估计适应需求,与智能体适应能力比较,超限时收紧动作集并激活屏蔽,实验验证该方法可减少安全违规,支持适应可行性原则。
AI 中文摘要
在非平稳性下确保强化学习中的安全性,需要确定学习系统能否在所需恢复范围内安全适应预测的环境变化。现有安全强化学习方法通常假设环境是平稳的,未明确将适应速度视为安全问题。然而,环境随时间演变时,延迟适应可能导致瞬时不安全行为。本文提出将调整速度作为非平稳强化学习的安全约束。核心思想是根据适应可行性定义安全性:当保持安全所需的适应超出学习系统校准的恢复能力时,未来状态或区域可能变得不安全。所提出的框架使用学习到的上下文表示和短视距上下文预测来估计适应需求,并将其与智能体可实现的适应能力进行比较。当预测的适应需求超过校准的恢复能力时,框架会主动收紧可允许动作集并激活动作级屏蔽,以在违规发生前减少不安全行为。在非平稳驾驶环境中的实验表明,所提出的方法主要减少了与上下文变化对齐的短视距窗口中的安全违规。消融研究进一步表明,屏蔽对于峰值和尾部风险抑制更为保守,而优化级调整在短视距切换条件违规方面提供了额外的减少。这些结果支持将适应可行性作为非平稳性下强化学习的实用安全原则,并表明主动干预可以在环境变化期间提高安全性。
英文摘要
Ensuring safety in reinforcement learning under nonstationarity requires determining whether a learning system can safely adapt to forecasted environmental change within the required recovery horizon. Existing safe reinforcement learning methods typically assume stationary environments and do not explicitly consider adaptation speed as a safety concern. However, when environments evolve over time, delayed adaptation may result in transient unsafe behavior. This paper proposes adjustment speed as a safety constraint for nonstationary reinforcement learning. The central idea is to define safety in terms of adaptation feasibility: future states or regions may become unsafe when the adaptation required to remain safe exceeds the learning system's calibrated recovery capacity. The proposed framework uses learned context representations and short-horizon context forecasts to estimate adaptation demand and compare it with the agent's achievable adaptation capacity. When predicted adaptation demand exceeds the calibrated recovery capacity, the framework proactively tightens the admissible action set and activates an action-level shield to reduce unsafe behavior before violations occur. Experiments in a nonstationary driving environment show that the proposed approach primarily reduces safety violations in short-horizon windows aligned with context changes. Ablation studies further show that shielding is more conservative for peak- and tail-risk suppression, while optimization-level adjustment provides additional reductions in short-horizon switch-conditioned violations. These results support adaptation feasibility as a practical safety principle for reinforcement learning under nonstationarity and demonstrate that proactive intervention can improve safety during periods of environmental change.
Comments15 pages, 5 figures, 2 tables. Preprint