赛道引导分层强化学习用于最小圈速规划的自动驾驶车辆漂移控制
Track-Guided Hierarchical Reinforcement Learning for Autonomous Vehicle Drifting with Minimum-Lap-Time Planning
AI总结:
该研究针对自动驾驶车辆漂移控制的双目标难题,提出TgRL分层强化学习框架,结合最优轨迹规划训练智能体,实现了兼顾控制性能与圈速缩短的漂移竞赛策略。
AI中文摘要:
在一级方程式赛车中,车手会在轮胎抓地力极限内优化走线以最小化圈速;而在拉力赛中,车手会在松散路面上故意打破抓地力进行漂移。这种操作能快速调整车辆姿态以利于弯道出口,最终缩短圈速。自主执行此类操作构成了一个复杂的双目标控制问题:既要稳定高度非线性的漂移动力学,又要严格最小化圈速。为解决这一挑战,本文开发了先进的最小圈速(Minimum-Lap-Time, MLT)漂移控制架构。本文提出了一种专门针对MLT漂移场景的规划-控制框架:首先,构建最优控制问题以生成MLT漂移规划轨迹,将其作为先验数据训练深度强化学习漂移控制器;考虑到漂移涉及极大侧滑角,直接学习极具挑战性,因此提出了赛道引导强化学习(Track-guided Reinforcement Learning, TgRL)漂移控制方法,以实现渐进式训练,依次从漂移控制策略、漂移弯道策略到综合漂移竞赛策略;奖励函数包含即时奖励项和源自MLT目标的最终奖励项。仿真结果表明,该框架使智能体能够学习漂移竞赛策略,不仅保证车辆运动控制性能,还能有效缩短圈速。
英文摘要:
In Formula 1, drivers optimize racing lines within tire grip limits to minimize lap times; however, in rally racing, drivers intentionally break traction to drift on loose surfaces. This maneuver rapidly aligns the vehicle for corner exits, ultimately reducing lap time. Autonomously executing such maneuvers formulates a complex dual-objective control problem: stabilizing highly nonlinear drift dynamics while strictly minimizing lap time. Addressing this challenge motivates the development of advanced Minimum-Lap-Time (MLT) drift control architectures. This paper proposes a planning-control framework specifically designed for MLT drifting scenario. First, we formulate an optimal control problem to generate a MLT drift planning trajectory, which is used as prior data to train a deep reinforcement learning drift controller. Given that drifting involves extremely large sideslip angles and is therefore challenging to learn directly, a Track-guided Reinforcement Learning (TgRL) drift control method is proposed to enable progressive training in a step-by-step manner, from drift control policy, to drift corner policy, and finally to a comprehensive drift race policy. The reward function incorporates both an instant reward term and an end reward term derived from the Minimum-Lap-Time objective. Simulation results demonstrate that the proposed framework enables the agent to learn a drift racing policy that not only ensures vehicle motion control performance but also effectively reduces lap time.