arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于分层优化和基于学习的控制的自主赛车安全超车

Safe Overtaking for Autonomous Racing Using Hierarchical Optimization and Learning-Based Control

Hassan Jardali, Kai Yin, Lantao Liu

arXiv 2607.13348首次发表:更新:

发表机构

Luddy School of Informatics, Computing, and Engineering, Indiana University; Expedia Group(印第安纳大学卢迪信息学、计算与工程学院; 亿客行集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究自主赛车安全超车问题,提出分层超车框架,高层MIQP解决超车拓扑选择,非线性Frenet框架MPC结合CBF约束保证安全,强化学习策略在线调整CBF衰减参数,提升性能与鲁棒性,实现安全与性能平衡。

AI 中文摘要

自主赛车超车需要在非线性车辆动力学和实时约束下平衡竞争性能与安全性。模型预测控制(MPC)与控制障碍函数(CBF)相结合,为验证安全集的前向不变性提供了一种有原则的机制。然而,常用的固定衰减离散时间CBF公式在交互式赛车场景中可能会变得过于保守,限制超车性能并需要在不同赛道条件下进行手动调整。本文提出了一种分层超车框架,明确将操纵级决策与安全认证轨迹控制分开,在保持安全性的同时减少保守性。高层混合整数二次规划(MIQP)通过选择可行的超车拓扑来解决组合超车侧选择问题,而非线性Frenet框架MPC通过嵌入离散时间CBF约束来强制执行车辆动力学和安全性。这种分解将操纵选择的组合复杂性与连续轨迹优化隔离开来。为了进一步减轻固定衰减障碍约束的敏感性,强化学习策略在线调整离散时间CBF衰减参数,实现安全裕度的上下文相关调制,而无需直接控制车辆输入。仿真和缩比硬件实验表明,没有单一的固定衰减参数能在所有赛道上实现一致的强大性能,而自适应策略实现了最高的总体成功率,并在无需逐赛道调整的情况下始终实现强大的安全 - 性能权衡,提高了对环境变化的鲁棒性,同时在标称操作中保持对安全约束的满足。

英文摘要

Autonomous racing overtaking requires balancing competitive performance with safety under nonlinear vehicle dynamics and real-time constraints. Model Predictive Control (MPC) combined with Control Barrier Functions (CBFs) provides a principled mechanism for certifying forward invariance of a safe set. However, commonly used fixed-decay discrete-time CBF formulations can become overly conservative in interactive racing scenarios, limiting overtaking performance and requiring manual tuning across track conditions. This paper proposes a hierarchical overtaking framework that explicitly separates maneuver-level decision making from safety-certified trajectory control, reducing conservatism while preserving safety. A high-level Mixed-Integer Quadratic Program (MIQP) resolves the combinatorial passing-side selection problem by selecting a feasible overtaking topology, while a nonlinear Frenet-frame MPC enforces vehicle dynamics and safety through embedded discrete-time CBF constraints. This decomposition isolates the combinatorial complexity of maneuver selection from the continuous trajectory optimization. To further mitigate the sensitivity of fixed-decay barrier constraints, a reinforcement learning policy adapts the discrete-time CBF decay parameter online, enabling context-dependent modulation of safety margins without directly controlling vehicle inputs. Simulation and scaled-hardware experiments show that no single fixed decay parameter achieves uniformly strong performance across tracks, whereas the adaptive strategy attains the highest aggregate success rate and consistently strong safety--performance trade-offs without per-track tuning, improving robustness to environment variation while maintaining safety constraint satisfaction in nominal operation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑