发表机构
Sharif University of Technology(谢里夫理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种将A3C强化学习算法与PID控制器结合的四旋翼控制方法,利用并行智能体动态优化PID参数,仿真表明其比A2C方法收敛更快、性能更优。
AI 中文摘要
四旋翼飞行器在许多应用中具有很大的实用性,但其非线性特性和对扰动的敏感性带来了巨大的控制挑战。基本的PID控制器通常不够复杂,难以应对这些复杂性。本文提出了一种将异步优势演员-评论家(A3C)算法与PID控制器相结合的控制系统,用于四旋翼的姿态和轨迹跟踪。A3C控制器使用并行智能体通过神经网络动态优化PID参数。针对互补系统的系统辨识模块对系统状态进行预测,以制定最优控制策略。将所提出的框架与标准演员-评论家(A2C)模型进行了比较。仿真结果验证了两种方法都能准确跟踪。然而,基于A3C的控制器在损失函数上收敛得更好,这通过奖励图和损失曲线得到证实,表明其参数优化更优。这表明基于A3C的方法在四旋翼控制中实现了更好的性能,有效地将强化学习与传统控制相结合,以实现更高的适应性。
英文摘要
Quadcopters offer great utility in many applications, but their nonlinear nature and disturbance sensitivity present great control challenges. Basic PID controllers are generally not sophisticated enough to cope with these complexities. This paper suggests a control system that integrates the Asynchronous Advantage Actor-Critic (A3C) algorithm with a PID controller for quadcopter attitude and trajectory tracking. The A3C controller uses parallel agents to optimize PID parameters dynamically using a neural network. A system identification module for the complementary system makes predictions about system states for optimal control policy. The proposed framework was compared with a standard actor-critic (A2C) model. Simulation results verify that they both track accurately. However, the A3C-based controller converges much more for the loss function, as evidenced by reward figures and loss curves, demonstrating better parameter optimization. This shows that A3C-based approach results in improved performance for the control of quadcopter, effectively integrating reinforcement learning and traditional control to achieve higher adaptability.