发表机构
Deakin University; Federation University Australia(迪肯大学; 澳大利亚联邦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对强化学习中分布外检测问题,提出OOD-RL-Bench框架,通过共享接口集成检测器与异常注入器,在LunarLander-v3环境评估其效用,揭示不同异常类型性能差异,还公开相关成果以便重现。
AI 中文摘要
可靠的强化学习(RL)智能体必须在传感器故障、动态干扰和缓慢的环境变化中保持操作完整性。分布外(OOD)条件的检测对于确定智能体的观察、转换或轨迹动态何时偏离其策略训练所依据的假设至关重要。当前的OOD检测基准通常评估图像分类器或静态低维数据集,未考虑RL轨迹中固有的复杂、与动作相关的时间结构。为解决这一差距,我们提出了OOD-RL-Bench,这是一个全面且可扩展的框架,旨在针对注入到RL轨迹中的异常类别评估OOD检测器。检测器和异常注入器通过共享接口和配置集成,无需修改核心基准循环即可评估新的评分方法和扰动族。我们在LunarLander-v3环境中使用深度Q网络策略评估了该框架的效用。我们使用匹配时间AUROC、匹配时间AUPRC、匹配时间误报率、检测延迟和分段起始指标评估了每个检测器在一系列异常类型上的性能。我们的分析揭示了不同异常类型之间存在显著的性能差异:几种方法能高精度识别观察扰动和状态切换,而观察延迟和动作条件动态即使将起始后异常分数与同一时间步的干净分数比较仍很困难。我们将框架、训练好的策略检查点和完整结果作为可重现的工件公开。
英文摘要
Reliable reinforcement learning (RL) agents must maintain operational integrity amidst sensor malfunctions, dynamic disturbances, and slow environmental shifts. The detection of out-of-distribution conditions is pivotal to determining when an agent's observations, transitions, or trajectory dynamics deviate from the assumptions underpinning its policy training. Current out-of-distribution (OOD) detection benchmarks typically evaluate image classifiers or static low-dimensional datasets, failing to account for the complex, action-dependent temporal structure inherent in RL trajectories. To address this gap, we present OOD-RL-Bench, a comprehensive and extensible framework designed to evaluate OOD detectors against categories of anomalies injected into RL trajectories. Detectors and anomaly injectors are integrated through shared interfaces and configuration, which allows new scoring methods and perturbation families to be evaluated without modification of the core benchmark loop. We evaluate the utility of the framework using a Deep Q-Network policy within the LunarLander-v3 environment. We assess the performance of each detector across a suite of anomaly types using matched-time AUROC, matched-time AUPRC, matched-time false-positive rate, detection delay, and segmented-onset metrics. Our analysis reveals significant performance variance across anomaly types: observation perturbations and regime switches are identified with high accuracy by several methods, while observation delay and action-conditioned dynamics remain difficult even when post-onset anomaly scores are compared against clean scores from the same timesteps. We make the framework, trained policy checkpoint, and complete results publicly available as a reproducible artefact.
Comments21 pages, 2 figures