AI 中文总结
研究GRPO对小语言模型用于四旋翼连续控制任务微调的情况,发现普通GRPO微调失败,两项消融实验虽防熵崩溃但无法学习,更换动作接口后训练收敛,描绘出帕累托前沿,该方法在三个模型上评估,还对比了经典基线并发现训练分布差距。
AI 中文摘要
我们研究了群体相对策略优化(GRPO)能否对用于模拟四旋翼连续控制任务的小语言模型进行微调。在我们的基准测试中,对Qwen-0.5B进行25Hz四旋翼速度控制的普通GRPO微调会退化为平凡的零动作:成功率为0%,熵在60步内从0.35降至0.03。两项消融实验——去除急动惩罚项和去除与预训练先验的KL锚点——都能防止熵崩溃,但都无法实现学习。当动作接口被替换为对PID预设的5路分类选择时,训练收敛。最终的控制器在训练过程中描绘出一条平滑度-可靠性帕累托前沿;报告了两个端点:在64步时,成功率为98.6%,急动为0.656m/s3;在256步时,成功率为100%,急动为1.103m/s3,或在匹配速度上限下为0.796。该方法在三个预训练语言模型上进行了评估。作为背景,重新调整的经典基线,即Ki = 0.30和vmax = 2.5的PID,在急动为0.736m/s3时达到相同的100%成功率。使用Crazyflie 2.1动力学的高保真模拟揭示了悬停区域训练分布差距。
英文摘要
We study whether Group Relative Policy Optimization (GRPO) can fine-tune small language models for simulated quadrotor continuous-control tasks. In our benchmark, vanilla GRPO fine-tuning of Qwen-0.5B for 25 Hz quadrotor velocity control collapses to the trivial zero action: 0 percent success rate, with entropy falling from 0.35 to 0.03 within 60 steps. Two ablations - removing the jerk-penalty term and removing the KL anchor to the pretrained prior - each prevent entropy collapse, yet neither enables learning. When the action interface is replaced by a 5-way categorical choice over PID presets, training converges. The resulting controller traces a smoothness-reliability Pareto frontier along training duration; both endpoints are reported: 98.6 percent success with 0.656 m/s3 jerk at 64 steps, and 100 percent success with 1.103 m/s3 jerk, or 0.796 under a matched velocity cap, at 256 steps. The recipe is evaluated across three pretrained language models. As context, a re-tuned classical baseline, PID with Ki = 0.30 and vmax = 2.5, reaches the same 100 percent success rate at jerk 0.736 m/s3. A high-fidelity simulation using Crazyflie 2.1 dynamics surfaces a hover-region training-distribution gap.