arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

能量感知路径跟踪:电动汽车强化学习与NMPC的对比分析

Energy-Aware Path Following: Comparative Analysis of Reinforcement Learning and NMPC for Electric Vehicles

Mohamed Sabaa, Mostafa Emam

arXiv 2610.08112首次发表:更新:

发表机构

Najran University; University of the Bundeswehr Munich(纳吉兰大学; 慕尼黑联邦国防军大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究对比了电动汽车路径跟踪中NMPC与PPO等四种控制器,利用VT-CPEM能量模型,验证了PPO策略可零样本迁移至新轨迹和动态模型,并兼顾路径偏差与能量回收。

AI 中文摘要

路径跟踪控制策略通常面临双目标优化难题:在最小化参考路径偏差的同时保持平滑的速度曲线。后者对于电动汽车(EVs)尤为重要,因为其有限的续航里程可以通过再生制动回收能量来延长,这一特性在文献中尚未得到充分研究。在本工作中,我们在一个共同的基于Frenet框架的运动学车辆模型下,对四种控制器进行了对比分析,并利用带有显式再生制动的验证能量模型(VT-CPEM)。在此,我们实现了以下控制器:非线性模型预测控制(NMPC)、近端策略优化(PPO)、增益调度Ackermann状态反馈基线(PID-SF)以及Stanley几何基线。为满足实时性要求,我们使用JIT编译的CasADi实现NMPC。此外,我们使用传统的直线和S曲线轨迹训练PPO,之后成功地将未修改的策略迁移到未见过的轨迹,包括:ISO 3888-1换道、弯道、随机生成的参数化样条以及±3°坡度道路。另外,该策略以零样本方式迁移到带有线性轮胎的动态单轨车辆模型,初始性能可接受,并在短暂微调后得到优化。由此,我们证明了我们的PPO可以轻松迁移到更全面的车辆模型。最后,我们对所开发的控制器进行了性能分析,并讨论了未来工作的想法。

英文摘要

Path-following control strategies typically follow the bi-objective optimization dilemma: minimizing deviations from a reference path while maintaining smooth speed profiles. The latter objective is especially relevant for Electric Vehicles (EVs), since their limited driving range can be extended by recovering energy through regenerative braking, a feature that has not yet been sufficiently studied in the literature. In this work, we perform a comparative analysis of four controllers under one common Frenet frame-based kinematic vehicle model, utilizing a validated energy model (VT-CPEM) with explicit regenerative braking. Herein, we implement the following controllers: Nonlinear Model Predictive Control (NMPC), Proximal Policy Optimization (PPO), gain-scheduled Ackermann state-feedback baseline (PID-SF), and a Stanley geometric baseline. To satisfy real-time requirements, we implement the NMPC using JIT-compiled CasADi. Moreover, we train the PPO using traditional straight and S-curve tracks, after which we successfully transfer the unmodified policy to unseen tracks, including: an ISO 3888-1 lane-change, a chicane, randomly-generated parameterized-splines, and a $\pm3^\circ$ graded road. In addition, the policy transfers to a dynamic single-track vehicle model with linear tires, zero-shot with an acceptable initial performance, which was optimized after brief fine-tuning. Thereby, we demonstrate that our PPO is readily transferable to more comprehensive vehicle models. We conclude with a performance analysis of developed controllers and discuss ideas for future work.

Comments20 pages, 12 figures, currently submitted for review at the journal (Robotics and Autonomous Systems) https://www.sciencedirect.com/journal/robotics-and-autonomous-systems

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑