DreamTest:用于深度强化学习智能体基于搜索测试的世界模型代理
DreamTest: World-Model Surrogates for Search-Based Testing of Deep Reinforcement Learning Agents
浏览论文内容
中文总结 AI 辅助
提出DreamTest世界模型代理,通过想象回合预测DRL智能体失败,指导搜索测试,在多个基准上显著提升失败预测精度和多样性。
中文摘要 AI 辅助
在信息物理系统中测试深度强化学习(DRL)智能体,旨在部署前发现各种故障,但每次执行可能代价高昂。代理辅助测试通过预测哪些测试配置可能失败来降低此成本。先前的代理将系统视为黑盒,直接预测通过或失败结果;我们则转而建模测试如何展开,并从想象的回合中估计失败。我们提出DreamTest,一种用于测试DRL智能体的世界模型代理。DreamTest调整循环状态空间模型,从智能体的训练日志中学习智能体行为和环境动态。给定候选配置,想象的滚动生成失败分数,指导搜索而无需在模拟器或真实系统中执行每个候选。我们在Parking、Humanoid和DonkeyCar上评估DreamTest的失败预测、测试生成和失败多样性。平均精确率-召回率曲线下面积(AUPRC)分别超过最强基线97%、12%和39%,在五个分布外测试集上的增益分别达到145%、29%和44%。在相同的模拟器验证预算下,最佳的“DreamTest + 搜索”组合平均发现29%、22%和79%更多的新颖失败。在k = 2-40的聚类中,DreamTest生成的失败几乎覆盖了所有k的最多行为聚类,表明DreamTest始终发现行为多样性的失败。
英文摘要
Testing deep reinforcement learning (DRL) agents in cyber-physical systems aims to uncover diverse failures before deployment, but each execution can be expensive. Surrogate-assisted testing reduces this cost by learning to predict which test configurations are likely to fail. Prior surrogates treat the system as a black box and predict pass or fail outcomes directly; we instead model how a test unfolds and estimate failure from an imagined episode. We introduce DreamTest, a world-model surrogate for testing DRL agents. DreamTest adapts a recurrent state-space model to learn agent behaviour and environment dynamics from the agent's training log. Given a candidate configuration, imagined rollouts produce a failure score that guides search without executing every candidate in a simulator or real system. We evaluate DreamTest for failure prediction, test generation, and failure diversity on Parking, Humanoid, and DonkeyCar. Mean area under the precision-recall curve (AUPRC) exceeds the strongest baseline by 97%, 12%, and 39%, respectively, and gains on five out-of-distribution test sets reach 145%, 29%, and 44%. Under the same simulator-validation budget, the best "DreamTest + search" combinations find 29%, 22%, and 79% more novel failures on average. Across clusterings with k = 2-40, failures generated with DreamTest cover the most behavioural clusters for almost all k, indicating that DreamTest consistently discovers behaviourally diverse failures.
发表机构
- Lero Research Centre(莱罗研究中心)
- University of Limerick(利默里克大学)
- University of Ottawa(渥太华大学)
机构由 AI 辅助整理,请以论文原文为准。