arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

混合交通环境下自动驾驶车辆控制的知识数据双驱动强化学习

Knowledge-Data-Dual-Driven Reinforcement Learning for Autonomous Vehicle Control in Mixed Traffic

Jie Fang, Wei Zheng, Mengyun Xu, Eui-Jin Kim

arXiv 2608.13878首次发表:更新:

AI 中文总结

针对混合交通下自动驾驶车辆控制的三个挑战,提出KDDRL算法,通过条件深度生成模型、双驱动范式与耦合模块实现优化,在仿真中表现更优。

AI 中文摘要

在混合交通场景中,自动驾驶车辆(AV)的决策面临三个相互关联的挑战:其一,强化学习(RL)模型中融入的基于物理的先验知识无法捕捉潜在的交互车辆意图和多样化的驾驶员行为,限制了主动推理能力;其二,周围车辆的突发操作会导致非平稳性,使得长尾安全事件未被充分探索;其三,由于连续跟驰和离散换道操作的时间尺度不同,混合动作空间会破坏统一的RL训练。为解决这些问题,本文提出知识数据双驱动强化学习(KDDRL):首先,条件深度生成模型合成感知意图的未来轨迹,将被动感知转化为主动预测状态;其次,知识数据双驱动范式基于这些预测状态运行,融合概率数据驱动的见解与物理约束,指导安全关键场景下的安全探索;最后,耦合模块将感知意图的轨迹和物理约束压缩为紧凑的共享嵌入,该统一表示能在保留互信息的同时,对连续跟驰和离散换道进行异步多时间尺度优化。在数据集校准的仿真中评估显示,KDDRL可有效处理意图不确定性,加快训练收敛速度,且在安全性、效率和舒适性方面优于传统基线方法。

英文摘要

In mixed traffic, decision-making for autonomous vehicles (AVs) confronts three interrelated challenges. First, physics-based priors incorporated into reinforcement learning (RL) models fail to capture latent interactive vehicle intentions and diverse driver behaviors, limiting the proactive reasoning capabilities. Second, abrupt maneuvers by surrounding vehicles cause non-stationarity, leaving long-tail safety events under-explored. Third, hybrid action spaces destabilize unified RL training due to the different temporal scales of continuous car-following and discrete lane-changing maneuvers. To address these issues, we propose Knowledge-Data Dual-driven Reinforcement Learning (KDDRL). First, a conditional deep generative model synthesizes intention-aware future trajectories, converting passive perception into proactive predictive states. Second, a knowledge-data dual-driven paradigm operates on these predictive states, fusing probabilistic data-driven insights with physical constraints to guide safe exploration through safety-critical scenarios. Third, a coupling module compresses both intention-aware trajectories and physical constraints into compact shared embeddings. This unified representation enables asynchronous multi-timescale optimization of continuous car-following and discrete lane-changing while preserving mutual information. Evaluations on dataset-calibrated simulations demonstrate that KDDRL effectively handles intention uncertainty, accelerates training convergence, and outperforms conventional baseline methods in terms of safety, efficiency, and comfort.

Comments16 pages, 17 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑