arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2510.24674cs.LGcs.AIcs.SYeess.SY

基于混合选项的安全驾驶学习

Learning to Drive Safely with Hybrid Options

  • STADIUS Center for Dynamical Systems, Signal Processing and Data Analytics, Department of Electrical Engineering (ESAT), KU Leuven(STADIUS 动力学系统、信号处理与数据分析中心,电气工程系(ESAT),KU 列文大学)

机构由 AI 辅助整理,请以论文原文为准。

Bram De Cooman, Johan Suykens

更新

AI总结:

本研究将选项框架适配于高速公路自动驾驶任务,定义带安全舒适约束的纵横向专用选项,提出多种分层控制方案及对应算法,其中混合选项策略兼具类人灵活性与可解释性,在多变交通下表现优于基线。

AI中文摘要:

在众多用于自动驾驶的深度强化学习方法中,仅有少数采用了选项(或称技能)框架。这一现象令人意外,因为该框架天然适用于通用的分层控制场景,尤其适配自动驾驶任务。因此,本研究将选项框架应用于高速公路自动驾驶任务并进行了针对性调整。具体而言,我们为纵向和横向机动定义了嵌入安全与舒适性约束的专用选项,借此可将先验领域知识融入学习过程,且更易约束习得的驾驶行为。我们提出了多种基于选项的分层控制方案,并结合前沿强化学习技术推导了实用算法。通过分别选择纵向和横向控制动作,所提出的基于组合与混合选项的策略具备与人类驾驶员同等的表达能力与灵活性,且比基于连续动作的经典策略更具可解释性。在所有研究的方法中,这类基于混合选项的灵活策略在多变交通条件下表现最佳,性能优于基于动作的基线策略。

英文摘要:

Out of the many deep reinforcement learning approaches for autonomous driving, only few make use of the options (or skills) framework. That is surprising, as this framework is naturally suited for hierarchical control applications in general, and autonomous driving tasks in specific. Therefore, in this work the options framework is applied and tailored to autonomous driving tasks on highways. More specifically, we define dedicated options for longitudinal and lateral manoeuvres with embedded safety and comfort constraints. This way, prior domain knowledge can be incorporated into the learning process and the learned driving behaviour can be constrained more easily. We propose several setups for hierarchical control with options and derive practical algorithms following state-of-the-art reinforcement learning techniques. By separately selecting actions for longitudinal and lateral control, the introduced policies over combined and hybrid options obtain the same expressiveness and flexibility that human drivers have, while being easier to interpret than classical policies over continuous actions. Of all the investigated approaches, these flexible policies over hybrid options perform the best under varying traffic conditions, outperforming the baseline policies over actions.

↑