发表机构
College of Computer and Data Science, Fuzhou University; Maynooth International College of Engineering, Fuzhou University; Institute of Automation, Chinese Academy of Sciences(福州大学计算机与数据科学学院; 福州大学梅努斯国际工程学院; 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对离线强化学习在自动驾驶中面临的分布偏移等挑战,提出含RHDif架构与3DICE范式的DiDrive框架,在CARLA基准的高密度交通场景中表现优于基线,为安全自动驾驶决策提供稳健路径。
AI 中文摘要
尽管扩散模型能有效捕捉自动驾驶的多模态行为先验,但离线强化学习(RL)策略仍易受分布偏移、重尾风险信号、分布外(OOD)动作生成及高维状态冗余的影响。为应对这些挑战,我们提出DiDrive,一种分布引导的离线扩散框架,包含两个协同组件:风险感知分层扩散(RHDif)架构与3DICE策略优化范式。在状态空间中,RHDif利用低层风险门控编码器与高层上下文调制器过滤环境冗余,聚焦安全关键威胁;在动作空间中,3DICE通过样本内校准引导、时空优化及基于集成的候选排序缓解OOD高估与梯度振荡。在CARLA基准上的评估显示,DiDrive优于IQL、CQL与Diffusion-QL等基线,尤其在含60辆车的复杂高密度交通场景中,其成功率达85%,平均奖励为4295.68,为安全自动驾驶决策提供了稳健路径。
英文摘要
While diffusion models effectively capture multimodal behavioral priors for autonomous driving, offline reinforcement learning (RL) policies remain susceptible to distribution shift, heavy-tailed risk signals, out-of-distribution (OOD) action generation, and high-dimensional state redundancy. To address these challenges, we propose DiDrive, a distribution-guided offline diffusion framework featuring two synergistic components: the Risk-Aware Hierarchical Diffusion (RHDif) architecture and the 3DICE policy optimization paradigm. In the state space, RHDif utilizes a low-level risk-gated encoder and a high-level contextual modulator to filter environmental redundancy and focus on safety-critical threats. In the action space, 3DICE mitigates OOD overestimation and gradient oscillation through in-sample calibrated guidance, spatiotemporal optimization, and ensemble-based candidate ranking. Evaluations on the CARLA benchmark demonstrate DiDrive's superiority over baselines like IQL, CQL, and Diffusion-QL, particularly in complex, high-density traffic scenarios with 60 vehicles, where it achieves an 85% success rate and a 4295.68 average reward, providing a robust pathway for safe autonomous driving decision-making.
Comments16 pages