发表机构
Singapore Management University; University of Alberta(新加坡管理大学; 阿尔伯塔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对深度强化学习智能体的最优性评估缺口,提出Delta框架,通过两阶段差异测试,结合离线RL算法,在5个环境中平均发现2518个最优性缺陷,效果优于基线方法50.2%。
AI 中文摘要
深度强化学习(DRL)在复杂决策问题中取得了显著成功,随着DRL系统越来越多地部署到实际应用中,确保其质量和可靠性至关重要。当前研究主要聚焦于检测安全关键型故障,却常常忽略策略最优性,这可能导致效率降低、用户信任度下降和经济损失。这种疏忽,加上最优性固有的“测试神谕问题”,在DRL系统的综合评估方面留下了重大缺口。为解决这一缺口,我们提出Delta(深度强化学习智能体的差异测试,Differential Testing for DRL Agents),一个新颖且全面的框架,可自动识别DRL智能体中的安全关键型故障和最优性缺陷。Delta采用两阶段方法:(1)安全测试,在此阶段评估被测智能体(Agent Under Test,AUT)的灾难性故障,同时收集其决策策略的数据;(2)最优性测试,利用前一阶段收集的数据通过离线强化学习训练一个挑战者智能体,随后将挑战者智能体与AUT进行差异测试;若挑战者获得更高的累积奖励,则表明AUT存在最优性问题。我们在5个环境中验证了Delta的有效性,研究了3种离线RL算法(BC、BCQ和CQL)在生成挑战者智能体时的有效性。实验结果表明,安全测试数据集对于训练有能力的DRL智能体具有价值,在Delta框架内,使用BCQ训练的挑战者智能体在识别最优性缺陷方面效果最佳。在5个环境中,Delta平均发现2518个最优性缺陷,比基线方法高出50.2%。
英文摘要
Deep Reinforcement Learning (DRL) has achieved significant success in complex decision-making problems. As DRL systems are increasingly deployed in real-world applications, ensuring their quality and reliability is paramount. Current works primarily focus on detecting safety-critical failures, often neglecting policy optimality, which can lead to reduced efficiency, user distrust, and economic losses. This oversight, compounded by the inherent "testing oracle problem" for optimality, leaves a significant gap in comprehensively evaluating DRL systems. To address this gap, we propose Delta (Differential Testing for DRL Agents), a novel and comprehensive framework that automatically identifies both safety-critical and optimality bugs in DRL agents. Delta employs a two-phase approach: (1) Safety Testing, where the Agent Under Test (AUT) is evaluated for catastrophic failures while collecting data from its decision-making policy, and (2) Optimality Testing, where this collected data from the prior phase is used to train a challenger agent via Offline Reinforcement Learning. Differential testing is then performed by comparing the challenger agent against the AUT; instances where the challenger achieves higher cumulative rewards indicate optimality issues in the AUT. We demonstrate Delta's effectiveness across five environments. We investigate the effectiveness of three offline RL algorithms (BC, BCQ, and CQL) in generating challenger agents. Experimental results demonstrate that safety testing datasets are valuable for training competent DRL agents. Challenger agents trained with BCQ proved most effective for identifying optimality issues within the framework of Delta. Across the five environments, Delta uncovered an average of 2,518 optimality issues, outperforming the baseline methods by 50.2%.