arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23506eess.SYcs.SY

稳定性感知的模型预测控制模仿学习用于自动驾驶车辆横向控制:精确Q损失与新颖训练流程

Stability-Aware Imitation Learning from Model Predictive Control for Autonomous Vehicle Lateral Control: Exact Q-Loss and a Novel Training Procedure

  • Ho Chi Minh City University of Technology (HCMUT), Vietnam National University Ho Chi Minh City (VNU-HCM)(胡志明市理工大学(HCMUT),越南国立大学胡志明市分校)
  • Hanoi University of Science and Technology (HUST)(河内科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Tien Dat Vu, Minh Quan Nguyen, Anh Tuan Vu, Thanh Tung Nguyen, Minh Doan

AI总结:

本文提出一种稳定性感知的模仿学习框架,通过精确Q损失和基于Lyapunov-IQC的认证裕度训练神经控制器逼近MPC策略,并在自动驾驶车辆横向控制实验中验证了闭环性能。

AI中文摘要:

本文开发了一个经过认证的模仿学习框架,用于用前馈神经控制器逼近模型预测控制(MPC)策略,并在自动驾驶车辆横向控制上进行了验证。通过在专家MPC问题中固定学习者的第一个转向动作并重新优化剩余时域,构建了精确的有限时域Q损失,从而衡量其下游最优控制后果,而不仅仅是逐点动作不匹配。神经策略被表示为线性分式变换(LFT)互连,激活非线性由扇形积分二次约束(IQCs)描述。结合二次Lyapunov条件,该表示产生基于Lyapunov-IQC矩阵最大特征值的可微认证裕度。训练期间通过对数障碍强制执行该裕度,而经过认证的数据聚合(DAgger)和安全投影将数据聚合回滚保持在认证策略集内。在带有AprilTag定位和实时转向的CAD参考自动驾驶车辆平台上的实验证明了所得到的闭环性能。

英文摘要:

This paper develops a certified imitation-learning framework for approximating model predictive control (MPC) policies with feedforward neural controllers and validates it on autonomous-vehicle lateral control. An exact finite-horizon Q-loss is constructed by fixing the learner's first steering action in the expert MPC problem and re-optimizing the remaining horizon, thereby measuring its downstream optimal-control consequence rather than only pointwise action mismatch. The neural policy is represented as a linear fractional transformation (LFT) interconnection with activation nonlinearities described by sector integral quadratic constraints (IQCs). Combined with a quadratic Lyapunov condition, this representation yields a differentiable certification margin based on the largest eigenvalue of the Lyapunov-IQC matrix. The margin is enforced during training through a logarithmic barrier, while certified Dataset Aggregation (DAgger) and safe projection keep data-aggregation rollouts within the certified policy set. Experiments on a CAD-referenced autonomous-vehicle platform with AprilTag localization and real-time steering demonstrate the resulting closed-loop performance.

↑