arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

固定时间弹性积分强化学习用于FDI攻击和扰动下输入受限未知非线性系统:一种数据驱动的容许热启动

Fixed-Time Resilient Integral Reinforcement Learning for Input-Constrained Unknown Nonlinear Systems Under FDI Attacks and Disturbances: A Data-Driven Admissible Warm Start

Tien Dat Vu, Minh Doan

arXiv 2609.14067首次发表:更新:

发表机构

Faculty of Mechanical Engineering, Ho Chi Minh City University of Technology (HCMUT), Vietnam National University Ho Chi Minh City (VNU-HCM)(越南胡志明市国家大学胡志明市理工大学机械工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对输入受限未知非线性系统在FDI攻击和扰动下的控制问题,提出一种固定时间弹性积分强化学习方法,利用数据驱动容许热启动,保证快速收敛和鲁棒性。

AI 中文摘要

本文针对在执行器限制、虚假数据注入攻击和外部扰动下运行的未知非线性系统,开发了一种弹性学习控制器。关键思想是直接从有限的轨迹数据中学习饱和安全策略,同时保证学习误差和闭环状态在独立于初始条件的统一固定时间内收敛到紧邻域。积分形式从可实施的学习律中消除了未知漂移,而存储的信息数据在在线激励消失后维持学习。为了减轻闭环对任意评论家初始化的敏感性,部署前数据(也可从回放堆栈中重用)通过有限维Koopman表示进行提升,以构建一个稳定的初始策略,其逆饱和策略映射提供了数据驱动的评论家权重热启动。所得控制器通过构造保持输入约束,并在持续攻击和扰动下保证实用的固定时间鲁棒性。所提出的学习和初始化架构进一步通过一个两连杆机器人稳定示例进行验证,结果表明在知情评论家初始化下,实现了快速状态恢复、有界评论家学习、可靠执行器约束满足以及改进的闭环行为。

英文摘要

This paper develops a resilient learning controller for unknown nonlinear systems operating under actuator limits, false-data-injection attacks, and external disturbances. The key idea is to learn a saturated secure policy directly from finite trajectory data while guaranteeing that both the learning error and the closed-loop state converge to compact neighborhoods within a uniform fixed time independent of initial conditions. An integral formulation removes the unknown drift from the implementable learning law, while stored informative data sustain learning after online excitation fades. To mitigate the closed-loop sensitivity to arbitrary critic initialization, pre-deployment data, which may also be reused from the replay stack, are lifted through a finite-dimensional Koopman representation to construct a stabilizing initial policy, whose inverse saturated-policy map provides a data-driven critic-weight warm start. The resulting controller preserves input constraints by construction and guarantees practical fixed-time robustness under persistent attacks and disturbances. The proposed learning and initialization architecture is further verified through a two-link robot stabilization example, where the results demonstrate rapid state recovery, bounded critic learning, reliable actuator-constraint satisfaction, and improved closed-loop behavior under informed critic initialization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑