增强深度强化学习的安全性:对抗攻击与防御综述
Enhancing Security in Deep Reinforcement Learning: A Comprehensive Survey on Adversarial Attacks and Defenses
- School of Computer and Information Engineering, Henan University(河南大学计算机与信息工程学院)
- Henan Industrial Technology Academy of Spatio-Temporal Big Data, Henan University(河南大学时空大数据工业技术研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文综述了深度强化学习面临的对抗攻击与防御方法,提出了攻击分类框架,总结了鲁棒性训练策略,并展望了未来研究方向。
AI中文摘要:
随着深度强化学习(DRL)技术在自动驾驶、智能制造和智能医疗等复杂领域的广泛应用,如何提高其在动态多变环境中的安全性和鲁棒性已成为当前研究的核心问题。尤其是在面对对抗攻击时,DRL可能遭受严重的性能下降,甚至做出潜在危险的决策,因此确保其在安全敏感场景中的稳定性至关重要。本文首先介绍了DRL的基本框架,并分析了在复杂多变环境中面临的主要安全挑战。此外,本文提出了一种基于扰动类型和攻击目标的对抗攻击分类框架,并详细综述了针对DRL的主流对抗攻击方法,包括扰动状态空间、动作空间、奖励函数和模型空间等多种攻击方法。为了有效应对这些攻击,本文系统总结了当前各种鲁棒性训练策略,包括对抗训练、竞争训练、鲁棒学习、对抗检测、防御蒸馏及其他相关防御技术,并讨论了这些方法在提高DRL鲁棒性方面的优缺点。最后,本文展望了DRL在对抗环境中的未来研究方向,强调了在提高泛化能力、降低计算复杂度以及增强可扩展性和可解释性方面的研究需求,旨在为研究人员提供有价值的参考和方向。
英文摘要:
With the wide application of deep reinforcement learning (DRL) techniques in complex fields such as autonomous driving, intelligent manufacturing, and smart healthcare, how to improve its security and robustness in dynamic and changeable environments has become a core issue in current research. Especially in the face of adversarial attacks, DRL may suffer serious performance degradation or even make potentially dangerous decisions, so it is crucial to ensure their stability in security-sensitive scenarios. In this paper, we first introduce the basic framework of DRL and analyze the main security challenges faced in complex and changing environments. In addition, this paper proposes an adversarial attack classification framework based on perturbation type and attack target and reviews the mainstream adversarial attack methods against DRL in detail, including various attack methods such as perturbation state space, action space, reward function and model space. To effectively counter the attacks, this paper systematically summarizes various current robustness training strategies, including adversarial training, competitive training, robust learning, adversarial detection, defense distillation and other related defense techniques, we also discuss the advantages and shortcomings of these methods in improving the robustness of DRL. Finally, this paper looks into the future research direction of DRL in adversarial environments, emphasizing the research needs in terms of improving generalization, reducing computational complexity, and enhancing scalability and explainability, aiming to provide valuable references and directions for researchers.