arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估强化学习智能体的模糊测试

Evaluating Fuzz Testing for Reinforcement Learning Agents

Zhibin Kang, Hanmo You, Dong Wang, Haiming Zheng, Junjie Chen

arXiv 2607.24577首次发表:更新:

发表机构

College of Intelligence and Computing, Tianjin University(天津大学智能与计算学部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对强化学习智能体模糊测试评估差异问题,通过在三种环境下统一配置对五种方法及随机测试进行基准测试,从多角度评估,揭示不同方法特点,表明模糊测试能提升智能体鲁棒性等,还给出可操作指导。

AI 中文摘要

强化学习智能体越来越多地应用于机器人技术、自动驾驶和无人机控制等安全关键领域,意外行为可能导致严重后果。模糊测试是探索强化学习智能体庞大状态空间并暴露崩溃的一种有前途的方法。尽管已经提出了许多强化学习模糊测试方法,但现有研究在评估设置、基线和指标方面往往存在差异。为填补这一空白,我们进行了首次全面实证研究,从有效性、多样性、效率和实际效用四个互补角度系统评估强化学习模糊测试方法。我们在三个复杂度递增的环境(MountainCar、BipedalWalker和CARLA)中,在统一配置下对五种先进方法和随机测试进行基准测试,并进一步评估检测到的崩溃对智能体鲁棒性改进和安全监控的下游有用性。结果揭示了几个关键见解,如MDPFuzz等面向吞吐量的方法在崩溃发现方面具有卓越的有效性和效率,而SeqDivFuzz等明确旨在鼓励探索的方法在发现多样崩溃行为方面表现出色。我们还表明,模糊测试生成的崩溃可以显著提高智能体鲁棒性,并通过强大的跨方法泛化实现准确的安全监控。此外,我们为研究人员和从业者提炼了可操作的指导,强调了结合互补模糊测试策略和采用多层次多样性分析以实现更全面和实际的强化学习测试的好处。

英文摘要

Reinforcement Learning (RL) agents are increasingly deployed in safety-critical domains such as robotics, autonomous driving, and drone control, where unexpected behaviors may lead to severe real-world consequences. Fuzz testing has recently emerged as a promising method for exploring the vast state spaces of RL agents and exposing crashes. Although numerous RL fuzzing methods have been proposed, existing studies often differ in evaluation settings, baselines, and metrics, making it difficult to draw reliable conclusions about their relative effectiveness and practical usefulness. To address this gap, we present the first comprehensive empirical study that systematically evaluates RL fuzzing methods from four complementary perspectives: effectiveness, diversity, efficiency, and practical utility. We benchmark five state-of-the-art methods alongside random testing under unified configurations across three environments of increasing complexity (MountainCar, BipedalWalker, and CARLA), and further assess the downstream usefulness of detected crashes for agent robustness improvement and safety monitoring. Our results reveal several key insights. For instance,throughput-oriented methods like MDPFuzz demonstrate superior effectiveness and efficiency in crash discovery, while methods explicitly designed to encourage exploration like SeqDivFuzz excel at uncovering diverse crash behaviors. We also show that fuzzing-generated crashes can meaningfully improve agent robustness and enable accurate safety monitoring with strong cross-method generalization. Beyond these empirical findings, we distill actionable guidance for both researchers and practitioners, highlighting the benefits of combining complementary fuzzing strategies and adopting multi-level diversity analysis to achieve more comprehensive and practical RL testing.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑