发表机构
Nanyang Technological University; University of California, Riverside; University of Connecticut; The Chinese University of Hong Kong, Shenzhen(南洋理工大学; 加利福尼亚大学河滨分校; 康涅狄格大学; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对深度强化学习的黑盒攻击,提出轨迹级可迁移对抗攻击,在CartPole-v1上跨模型和算法设置中优于每步攻击和随机噪声。
AI 中文摘要
大多数针对深度强化学习(DRL)的对抗攻击假设对受害策略具有白盒访问权限,而这在实际中很少成立。本文研究基于迁移的深度强化学习黑盒攻击:攻击者在白盒代理智能体上构造观测扰动,并将其输入到未知的受害智能体。我们将攻击形式化为在每步扰动预算下的回报最小化问题。我们首先表明,将可迁移的图像分类攻击(FGSM、MI-FGSM和NI-FGSM)与每步目标结合,所产生的扰动虽然可迁移,但强度不超过相同预算下的随机噪声。随后,我们提出一种轨迹级攻击,通过环境可微模型和温度平滑的代理策略,在递推时域上优化扰动序列,并使用相同的优化器。在CartPole-v1上,使用十个DQN和DDQN智能体以及100个代理-受害智能体对,轨迹级攻击在白盒、跨模型和跨算法设置中均优于每步攻击和随机噪声。
英文摘要
Most adversarial attacks on deep reinforcement learning (DRL) assume white-box access to the victim policy, which rarely holds in practice. This paper studies transfer-based black-box attacks on DRL: the attacker crafts observation perturbations on a white-box surrogate agent and feeds them to an unknown victim. We formulate the attack as return minimization under a per-step perturbation budget. We first show that transplanting transferable image-classification attacks (FGSM, MI-FGSM, and NI-FGSM) with a per-step objective yields perturbations that transfer but are no stronger than random noise of the same budget. We then propose a trajectory-level attack that optimizes a sequence of perturbations over a receding horizon through a differentiable model of the environment and a temperature-smoothed surrogate policy, with the same optimizers. On CartPole-v1 with ten DQN and DDQN agents and 100 surrogate--victim pairs, the trajectory-level attack outperforms per-step attacks and random noise in the white-box, cross-model, and cross-algorithm settings.