面向激进四旋翼飞行的电池感知强化学习
Battery-Aware Reinforcement Learning for Aggressive Quadrotor Flight
浏览论文内容
中文总结 AI 辅助
本研究提出电池感知强化学习策略,利用额外推力提升激进四旋翼飞行性能,在Crazyflie无人机上降低跟踪误差和比赛时间,并验证电压输入的有效性。
中文摘要 AI 辅助
无人机竞速和追逃等敏捷飞行任务需要强大的加速度和精确的转弯,但可用推力会随着电池放电和负载下电压下降而变化。保守的命令限制使这种变化更容易容忍,但代价是未使用的性能。我们研究了学习型控制器如何在保留飞行控制器的电压补偿和速率控制的同时,利用这些额外的推力。我们的训练模拟器将识别的负载瞬态电池模型与转子动力学和固件饱和相结合。前馈策略在训练和部署期间都接收滤波后的电压。受控消融实验区分了更大推力命令范围与电压信息带来的益处。在一架38克的Crazyflie Brushless无人机上,与标准权限的强化学习相比,所得策略在3.84米/秒的速度下将圆形轨迹跟踪误差降低了49%,同时保持了简单任务的精度。平均20圈比赛时间从106.22秒减少到95.24秒。与具有相同增加权限的电压盲策略相比,硬件误差和比赛时间分别降低了15.3%和4.5%。在仿真中,用不同电池条件下的记录替换策略的电压输入会恶化硬圆形轨迹跟踪,而在竞速中影响较小且依赖于电压。总之,这些结果表明,在激进的习得飞行中,简单的电压输入在何处补充了现有的执行器补偿。
英文摘要
Agile flight tasks such as drone racing and pursuit-evasion require strong acceleration and precise turns, but the available thrust changes as the battery discharges and voltage drops under load. Conservative command limits make this variation easier to tolerate, at the cost of unused performance. We investigate how learned controllers can use that additional thrust while retaining the flight controller's voltage compensation and rate control. Our training simulator couples an identified load-transient battery model to rotor dynamics and firmware saturation. The feedforward policy receives filtered voltage during both training and deployment. Controlled ablations distinguish the benefit of a larger thrust-command range from that of voltage information. On a 38 g Crazyflie Brushless, the resulting policy reduces circle tracking error by 49% relative to stock-authority RL at 3.84 m/s, while preserving easy-task precision. Mean 20-lap race time decreases from 106.22 s to 95.24 s. Compared with a voltage-blind policy with the same increased authority, hardware error and race time are lower by 15.3% and 4.5%, respectively. In simulation, replacing the policy's voltage input with a recording from a different battery condition worsens hard-circle tracking, with a smaller, voltage-dependent effect in racing. Together, these results show where a simple voltage input complements existing actuator compensation in aggressive learned flight.
发表机构
- Royal Institute of Technology (KTH)(皇家理工学院(KTH))
机构由 AI 辅助整理,请以论文原文为准。