发表机构
Ho Chi Minh City University of Technology (HCMUT), Vietnam National University Ho Chi Minh City (VNU-HCM)(胡志明市理工大学,越南国立大学胡志明市分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于逆最优设计的预定时积分强化学习框架,用于未知非线性系统最优控制,通过RBF网络近似漂移项与值函数,实现评论器权重预定时收敛,仿真验证了方法有效性。
AI 中文摘要
本文针对未知非线性系统的最优控制问题,提出了一种新的预定时积分强化学习框架。首先采用径向基函数(RBF)神经网络近似未知漂移项,并结合数据驱动的在线辨识律更新对应的神经权重。当辨识器收敛到真实动力学的足够小邻域后,将学习到的模型融入积分强化学习(IRL)问题。与传统基于强化学习的最优控制不同,本文将期望收敛时间直接引入控制目标:设计者指定李雅普诺夫函数及其规定的衰减行为,随后利用逆最优控制构造相容的运行代价,其最优策略继承预定时镇定特性。采用第二个RBF神经网络近似值函数,并开发新的评论器更新律以确保评论器权重的预定时收敛。有限的有效学习数据存储在回放缓冲器中,并在评论器更新期间重复使用,从而避免闭环运行全程需要持续激励。理论分析证明,评论器权重误差在分配的学习时域内进入规定的残差集,而闭环状态在设计者指定的总时限内到达原点的小邻域。对未知非线性系统的数值仿真验证了漂移项的精确重构、评论器的预定时学习以及闭环收敛性。
英文摘要
This paper develops a new predefined-time integral reinforcement learning framework for optimal control of unknown nonlinear systems. The unknown drift is first approximated by a radial basis function (RBF) neural network, together with a data-driven online identification law for updating the corresponding neural weights. After the identifier converges to a sufficiently small neighborhood of the true dynamics, the learned model is incorporated into the integral reinforcement learning (IRL) problem. Unlike conventional reinforcement-learning-based optimal control, the desired convergence time is introduced directly into the control objective: a Lyapunov function and its prescribed decay behavior are specified by the designer, and inverse-optimal control is then used to construct a compatible running cost whose optimal policy inherits the predefined-time stabilization property. The value function is approximated by a second RBF neural network, and a new critic update law is developed to impose predefined-time convergence on the critic weights. Finite informative learning data are stored in a replay buffer and reused during the critic update, thereby avoiding the need for persistent excitation throughout the closed-loop operation. Theoretical analysis proves that the critic-weight error enters a prescribed residual set within the allocated learning horizon, while the closed-loop state reaches a small neighborhood of the origin within the overall designer-specified deadline. Numerical simulations on an unknown nonlinear system verify accurate drift reconstruction, predefined-time critic learning, and closed-loop convergence.