arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种改进的动车组牵引双整流器深度强化学习控制策略

An Improved Deep Reinforcement Learning Control Strategy for Traction Dual Rectifiers in EMUs

Zhigang Liu, Mingwei Tang, Xiangyu Meng, Hui Wang, Qiao Zhang, Haoyu Wang, Mengru Li

arXiv 2607.09276首次发表:更新:

AI 中文总结

研究动车组牵引双整流器控制策略,针对现有PI控制问题,提出基于深度强化学习的新策略,改进双延迟深度确定性策略梯度算法,添加奖励塑造并结合优先经验回放,经仿真和稳定性分析验证,该策略能有效应用于多工况动车组。

AI 中文摘要

由于CRH5高速列车脉冲整流器中采用基于PI的d q电流解耦,PI参数直接影响牵引系统控制性能。线性化控制在参考轨迹变化或模型失配时存在问题,非线性控制存在抖动和稳态精度差的问题。本文提出用单个智能体取代d q电流解耦控制中的所有PI的新控制策略。基于深度强化学习(DRL)的方法可避免线性化和非线性控制的缺点并确保中间直流电压稳定。但动车组在不同工况切换时,牵引双整流器使用的双延迟深度确定性策略梯度(TD3)算法控制效果不佳。针对此问题,添加奖励塑造(RS)重新设计非线性奖励函数,可与优先经验回放(PER)结合提高情节奖励收敛速度。仿真结果表明改进控制策略可有效应用于多工况动车组。最后用李雅普诺夫第二方法进行稳定性分析,硬件在环(HIL)仿真平台验证结果表明DRL控制效果良好。

英文摘要

Due to the use of PI-based d q current decoupling in the pulse rectifier of CRH5 high-speed trains, the PI parameters directly affect the traction system's control performance. Linearized control may have issues with reference trajectory changes or model mismatches, leading to a decrease in system performance, while nonlinear control may have problems with jitter and poor steady-state accuracy. This paper proposes a new control strategy that replaces all PI in the d q current decoupling control with a single intelligent agent. This method based on Deep Reinforcement Learning (DRL) can avoid various drawbacks of linearization and nonlinear control and ensure the stability of intermediate DC voltage. However, when EMUs are in different working conditions and switching, the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm used in traction dual rectifiers does not have a good control effect. Focusing on the issue, Reward Shaping (RS) is added to re-design a nonlinear reward function, which can be combined with Prioritized Experience Replay (PER) to increase the convergence speed of the episode reward. The simulation results show that the improved control strategy can be effectively applied to EMUs working in multiple conditions. Finally, the stability analysis is carried out using Lyapunov's second method and the verification results of the hardware-in-the-loop (HIL) simulation platform show that the DRL control has a good effect.

Comments19 pages. Accepted manuscript

Journal refIEEE Transactions on Transportation Electrification, pp. 1-1, 2026

DOI:10.1109/TTE.2026.3660375

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑