AI 中文总结
研究针对四旋翼无人机自主控制挑战,采用基于物理感知的端到端深度强化学习方法,结合特定奖励和执行器动力学模型,评估四种算法,发现SAC和TD3表现优,强调建模重要性并提供基准。
AI 中文摘要
无人机,尤其是四旋翼无人机,由于其欠驱动动力学特性,在自主控制方面面临独特挑战:四个控制输入要控制六个自由度。本文研究一种基于物理感知的端到端深度强化学习方法,直接作用于低级机体输入(总推力和机体扭矩$(T,\tau_x,\tau_y,\tau_z)$),并通过高保真Simulink环境闭环。模拟器集成12状态刚体模型,有基于系数矩阵伪逆的动作到转速分配及每个电机的一阶执行器动力学。通过特定奖励平衡目标达成和稳定性。评估了四种深度强化学习算法在两个阶段的表现。结果表明SAC和TD3稳定性和探索效率更高,PPO样本效率较低。该研究突出了对执行器滞后和气动矩建模对稳定低级控制的重要性,并为四旋翼深度强化学习提供了可重现基准。
英文摘要
Unmanned aerial vehicles (UAVs), particularly quadcopters, present unique challenges for autonomous control due to their underactuated dynamics: only four available control inputs must govern six degrees of freedom. This paper investigates a physics-aware, end-to-end deep reinforcement learning (DRL) approach that acts directly on low-level body inputs, total thrust and body torques $(T, τ_x, τ_y, τ_z)$, and closes the loop through a high-fidelity Simulink environment. Our simulator integrates a 12-state rigid-body model (MATLAB Level-2 S-Function) with (i) an Action2RPM allocation based on the Moore-Penrose pseudo-inverse of a coefficient matrix derived from thrust and drag terms, and (ii) first-order actuator dynamics for each motor (time constant $T_m = 0.076$ s), including rotor gyroscopic coupling. A shaped reward balances goal-reaching and stability using an exponential position well, attitude penalties, and quadratic velocity costs. Four DRL algorithms, DDPG, TD3, PPO, and SAC, are evaluated in two stages: (S1) thrust-only hover and (S2) hover with pitch torque and a translated goal. Results show that SAC and TD3 achieve superior stability and exploration efficiency, while PPO is less sample-efficient. The study highlights the significance of modeling actuator lags and aerodynamic moments for stable low-level control and provides a reproducible benchmark for quadcopter DRL.
Comments8 pages, 5 figures, 2 tables. Presented at the Aeronautical and Astronautical Society of the Republic of China (AASRC) Conference, Tamsui, Taiwan, November 15, 2025, Paper No. 1075. Received an Honorable Mention for the Best Paper Award