发表机构
Tongji University; The University of Hong Kong(同济大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出RUDC框架,融合风险敏感分布强化学习与不确定性量化,并采用自适应安全校正机制,在无信号交叉口实现安全、高效且鲁棒的自动驾驶决策。
AI 中文摘要
强化学习(RL)在自动驾驶决策中已展现出巨大潜力。然而,其在城市自动驾驶中的应用,尤其是在高交互性的无信号交叉口场景中,仍面临挑战,因为学习到的策略可能难以在复杂交通状况下同时保持安全性和鲁棒决策。传统的安全过滤方法通常采用固定的保守约束,这虽然可能提高安全性,但代价是过度干预和交通效率下降。为解决这些局限,我们提出了一种面向安全鲁棒自动驾驶的风险敏感与不确定性感知决策与控制(RUDC)框架。RUDC将风险敏感分布强化学习与基于集成的策略不确定性量化相结合,共同考虑回报分布中的尾部风险和所学策略中的不确定性。一种基于不确定性感知的高阶控制屏障函数(HOCBF)的安全校正机制根据策略不确定性自适应调整约束严格程度,同时一个可学习的残差预测器补偿CBF模型失配和离散化误差。在无信号交叉口的广泛仿真表明,RUDC在安全性、效率和鲁棒性之间取得了良好平衡,在标称场景及具有挑战性的分布外(OOD)和长尾场景下均优于具有代表性的安全RL基线,同时满足实时性要求。
英文摘要
Reinforcement learning (RL) has demonstrated considerable potential for autonomous driving decision-making. However, its deployment in urban autonomous driving, particularly at highly interactive unsignalized intersections, remains challenging, as learned policies may struggle to maintain both safety and robust decision-making in complex traffic situations. Conventional safety-filtering approaches typically employ fixed conservative constraints, which may improve safety at the cost of excessive intervention and degraded traffic efficiency. To address these limitations, we propose a Risk-sensitive and Uncertainty-aware Decision-making and Control (RUDC) framework for safe and robust autonomous driving. RUDC couples risk-sensitive distributional RL with ensemble-based policy uncertainty quantification, jointly accounting for tail risks in return distributions and uncertainty in learned policies. An uncertainty-aware high-order control barrier function (HOCBF)-based safety correction mechanism adaptively adjusts constraint strictness according to policy uncertainty, while a learnable residual predictor compensates for CBF model mismatches and discretization errors. Extensive simulations at unsignalized intersections demonstrate that RUDC achieves a favorable balance among safety, efficiency, and robustness, outperforming representative safe RL baselines under both nominal and challenging OOD and long-tail scenarios while satisfying real-time requirements.
Comments14 pages, 8 figures, 5 tables