arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CORE-RL:黑盒强化学习策略的置信度导向可靠性评估

CORE-RL: Confidence-Oriented Reliability Evaluation of Black-Box Reinforcement Learning Policies

Santhosh GS, Ananya Ravi, Devika Jay, Abhishek Sarkar, Perepu Satheesh Kumar, Saurav Prakash, Kaushik Dey, Balaraman Ravindran

arXiv 2610.04418首次发表:更新:

发表机构

Center for Responsible AI; IIT Madras; Ericsson Research(负责任人工智能中心; 印度理工学院马德拉斯分校; 爱立信研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对黑盒强化学习策略的可靠性评估难题,提出CORE-RL框架,通过统一可靠性度量与统计置信界,实现自动拒绝不合规策略并绘制安全操作域。

AI 中文摘要

在关键领域部署强化学习(RL)智能体之前,必须通过一个流程来评估RL智能体与复杂多目标规范的符合程度以及在真实世界环境漂移下的鲁棒性。然而,为了保护知识产权,RL智能体可能以不透明的可执行文件或远程API形式交付用于评估,这使得基于策略内部结构的传统评估技术变得不可行。为解决这一差距,本文提出了CORE-RL:黑盒RL策略的置信度导向可靠性评估。CORE-RL流程引入了一个统一可靠性度量,该度量正式整合了早期任务终止和安全约束违反,防止不安全策略通过过早终止回合来掩盖失败。通过将策略置于包含感知噪声、执行噪声和环境动力学变化的噪声认证包络中,该流程计算统一可靠性度量的有限样本Clopper-Pearson界以及奖励和安全成本的Hoeffding下界。随后,该流程定义安全操作设计域,以报告安全性和预期性能的高置信度证书。在连续控制任务上的实验表明,CORE-RL流程能够自动拒绝不符合要求的策略,并绘制安全感知策略的安全操作设计域(ODD)。因此,CORE-RL提供了一个评估框架,为安全评估、比较和部署黑盒RL解决方案提供了必要的定量、透明且可复现的统计依据。

英文摘要

The deployment of Reinforcement Learning (RL) agents in critical domains must be preceded with a pipeline to evaluate the alignment of the RL agent with complex multi-objective specifications and robustness under real-world environmental drift. However, to protect intellectual property, the RL agent may be delivered for evaluation as opaque executable or remote API, which makes traditional evaluation techniques based on the internals of the policies infeasible. To address this gap, CORE-RL: Confidence-Oriented Reliability Evaluation of black box RL policy is proposed in this paper. The CORE-RL pipeline introduces a Unified Reliability Metric that formally integrates early task termination and safety constraint violations, preventing unsafe policies from masking failures through premature episode halts. By subjecting the policy to a noise certification envelope of perceptual noise, actuation noise and change in environment dynamics, the pipeline computes the finite-sample Clopper-Pearson bounds on unified reliability metric and Hoeffdings' lower bound on reward and safety cost. The pipeline then defines safe operational design domain to report high-confidence certificates for safety and expected performance. Experiments on continuous control tasks demonstrate the CORE-RL pipeline's ability to automatically reject non-compliant policies and map the safe Operational Design Domain (ODD) of safety-aware policies. Thus CORE-RL provides an evaluation framework towards a quantitative, transparent and reproducible, statistical rationale necessary to safely evaluate, compare, and deploy black box RL solutions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑