AI 中文总结
本文提出结合SAC与MPC的混合控制架构,在多实验中验证其兼具SAC跟踪性能与MPC安全约束,可减弱分布偏移下的失效,还讨论了相关局限性与改进方向。
AI 中文摘要
联网自动化车辆需要横向控制器在模型误差和传感器噪声下同时具备高精度、低负荷和安全性。模型预测控制(Model Predictive Control, MPC)等模块化控制器可解释且感知约束,但依赖精确模型和手动调优的权重;端到端学习策略,尤其是连续动作深度强化学习,适应性强且无需手动设计控制律,但无内在安全保证且可解释性有限。本文提出一种混合架构,将端到端软演员-评论家(Soft Actor-Critic, SAC)策略与受约束线性MPC结合为单一转向指令,利用MPC的第一步最优值作为基于模型的锚点,通过单一单调混合系数在两种范式间插值。该架构在名义工况、单轴鲁棒性和多初始条件集成实验中,针对线性化横向自行车模型,与PID基线、调优后的线性MPC及独立SAC策略进行评估。混合架构保留了独立SAC的跟踪性能,同时保持在MPC的执行器包络内,并为每个转向指令保留确定性的基于模型的贡献。该架构通过构造提供执行器包络保证,但未建立递归可行性或终端不变性,且闭式混合无法阻止训练分布边界下的所有极端情况分歧。极端情况分析表明,该混合方案可减弱但无法阻止分布偏移下的失效,因此提出一种感知连通性的扩展方案,其中混合系数由车万物联网(Vehicle-to-Everything, V2X)信号调度以恢复基于模型的控制权。本文还讨论了局限性及通往约束二次规划(constrained-QP)预测安全滤波器的路径。
英文摘要
Connected and automated vehicles demand lateral controllers that are simultaneously accurate, low-effort, and safe under model error and sensor noise. Modular controllers such as model predictive control (MPC) are interpretable and constraint-aware but rely on accurate models and hand-tuned weights. End-to-end learned policies, in particular continuous-action deep reinforcement learning, are adaptable and require no hand-designed control law, but offer no intrinsic safety guarantees and limited interpretability. This paper presents a hybrid architecture that combines an end-to-end Soft Actor-Critic (SAC) policy with a constrained linear MPC into a single steering command, using the MPC's first-step optimum as the model-based anchor and a single monotone blending coefficient that interpolates between the two paradigms. The architecture is evaluated on a linearized lateral bicycle model against a PID baseline, a tuned linear MPC, and a stand-alone SAC policy, across nominal, single-axis robustness, and multi-initial-condition ensemble experiments. The hybrid retains the tracking quality of stand-alone SAC while remaining inside the MPC's actuator envelope and preserving a deterministic, model-based contribution to every steering command. The architecture provides an actuator-envelope guarantee by construction but does not establish recursive feasibility or terminal invariance, and the closed-form blend does not prevent all corner-case divergences at the boundary of the training distribution. A corner-case analysis shows that the blend attenuates but cannot prevent failure under distribution shift, motivating a connectivity-aware extension in which the blending coefficient is scheduled by vehicle-to-everything (V2X) signals to restore model-based authority. Limitations and a path toward a constrained-QP predictive safety filter are discussed.