arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

反转自触发控制:用于稀疏拒绝服务攻击的对抗强化学习

Inverting Self-Triggered Control: Adversarial Reinforcement Learning for Sparse Denial-of-Service Attacks

Adam Haroon, Erick J. Rodríguez-Seda, Tristan Schuler, Cody Fleming

arXiv 2609.12016首次发表:更新:

发表机构

Iowa State University; United States Naval Academy; U.S. Naval Research Laboratory(爱荷华州立大学; 美国海军学院; 美国海军研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将自触发控制反转,提出对抗强化学习生成最稀疏的拒绝服务攻击调度,以李雅普诺夫递增条件破坏闭环稳定性,实验证明其在多种环境下以100%成功率优于基线。

AI 中文摘要

自触发强化学习控制(RL-STC)在运行时保证(RTA)覆盖下学习最稀疏的控制调度,该调度保持李雅普诺夫递减稳定性。我们将其反转:一个对抗性强化学习智能体学习最稀疏的干扰或拒绝服务(DoS)调度,使闭环失稳,其李雅普诺夫递增可接受谓词镜像了防御方的安全证书。我们证明了一个植物属性下界,即对于满足李雅普诺夫契约的自触发控制器(STC),立即保持最后介质访问控制(hold-last medium-access-control)对手迫使其崩溃所需的最小干扰次数,并作为推论恢复了先前计数预算DoS调度中连续分组最优性的证书级类比。这将DoS调度计数预算分析从周期性和线性时不变控制器扩展到STC控制器。在实验上,我们在Pendulum、CartPole和Quadrotor2D上针对每个植物的四个固定防御者(一个线性二次调节器(LQR)和三个RL-STC)进行训练。学习到的对手是唯一能在每个植物上以100%概率使每个防御者崩溃的对手:贪婪对手在42%的回合中错过Quadrotor2D LQR,而周期性对手在97%的回合中错过Pendulum LQR。在每次失败的干扰时间上,它比基线最多提升2.8倍,并在Quadrotor2D LQR上显示出最宽的绝对差距。鲁棒性消融实验表明,超过初始状态幅度的高斯观测噪声和仅位置观测都保持了100%的失败率,并使学习到的对手在每次失败的干扰时间上严格领先于两个基线。

英文摘要

Self-triggered reinforcement learning control (RL-STC) learns the sparsest control schedule that preserves Lyapunov-decreasing stability under a Run-Time Assurance (RTA) override. We invert this: an adversarial RL agent learns the sparsest jamming or Denial-of-Service (DoS) schedule that destabilizes the closed loop, with a Lyapunov-increase admissibility predicate mirroring the defender's safety certificate. We prove a plant-property lower bound on the minimum jam count required for an immediate hold-last medium-access-control adversary to force a crash against a self-triggered controller (STC) satisfying a Lyapunov contract, and recover a certificate-level analog of the consecutive-grouping optimality of prior count-budget DoS scheduling as a corollary. This extends the DoS-scheduling count-budget analysis from periodic and linear-time-invariant to STC controllers. Empirically, we train against four fixed defenders per plant (one Linear Quadratic Regulator (LQR) and three RL-STC) on Pendulum, CartPole, and Quadrotor2D. The learned adversary is the only adversary that crashes every defender on every plant at $100\%$: greedy misses Quadrotor2D LQR on $42\%$ of episodes and periodic misses Pendulum LQR on $97\%$. On jam-time-per-failure it beats baselines by up to $2.8\times$, and shows its widest absolute margin on Quadrotor2D LQR. Robustness ablations show that Gaussian observation noise exceeding the initial-state magnitude and position-only observation both preserve $100\%$ failure rate and keep the learned adversary strictly ahead of both baselines on jam-time-per-failure.

Comments8 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑