利用被动动力学实现弹簧腿四旋翼飞行器节能目标跳跃的学习方法
Learning to Exploit Passive Dynamics for Energy-Efficient Target Hopping of a Spring-Legged Quadcopter
浏览论文内容
中文总结 AI 辅助
针对弹簧腿四旋翼飞行器,提出基于PPO的直接状态到电机控制策略,通过能量与效率塑形奖励,相比PID控制降低电功率30.7%和推力49.8%,实现节能目标跳跃。
中文摘要 AI 辅助
将空中推力与弹簧加载跳跃相结合,使得单腿四旋翼飞行器在复杂地形上移动具有前景,但启发式比例-积分-微分(PID)调参限制了主动推力与被动接触动力学之间的协调。我们提出了一种直接从估计状态到电机的近端策略优化(PPO)策略,该策略无需显式的跳跃状态机或底层姿态PID即可控制四个电机。其奖励结合了用于质量归一化垂直能量跟踪和顶点状态锚定的能量流形塑形,以及效率塑形,后者利用基于历史信息的功率估计器来惩罚一般功率使用、施加额外的空中功率成本,并惩罚空中接近静止状态。在代表性硬件运行中,与调参后的基于PID的控制栈相比,基于PPO的控制栈将周期平均实测电功率降低了30.7%,平均总归一化推力降低了49.8%,同时保持了可重复的指令高度跳跃和更集中的着陆。这些观察结果与被动动力学的更好利用以及实测电力需求的降低相一致。
英文摘要
Combining aerial thrust with spring-loaded hopping makes monopedal quadcopters promising for locomotion over complex terrain, but heuristic proportional-integral-derivative (PID) tuning limits coordination between active thrust and passive contact dynamics. We present a direct estimated-state-to-motor Proximal Policy Optimization (PPO) policy that commands four motors without an explicit hopping state machine or low-level attitude PID. Its reward combines Energy-Manifold Shaping for mass-normalized vertical-energy tracking and apex-state anchoring with Efficiency Shaping, which uses a history-aware power estimator to penalize general power use, impose an additional airborne-power cost, and penalize airborne near-stationarity. In representative hardware runs, the PPO-based control stack reduced cycle-averaged measured electrical power by 30.7% and mean total normalized thrust by 49.8% relative to the tuned PID-based control stack, while retaining repeatable commanded-height hopping and more concentrated landings. These observations are consistent with improved use of passive dynamics and reduced measured electrical demand.
发表机构
- Guangdong Technion - Israel Institute of Technology(广东以色列理工学院)
- Technion-Israel Institute of Technology(以色列理工学院)
机构由 AI 辅助整理,请以论文原文为准。