用于在线强化学习安全自动驾驶的威胁引导、策略感知场景扰动
Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
本文提出TPSP方法,通过策略感知场景编码器与威胁引导优化策略,提升在线RL安全自动驾驶的学习效率,在NAVSIM v2数据集上用约400万公里模拟数据实现出色安全性能。
中文摘要 AI 辅助
强化学习(RL)在自动驾驶领域展现出良好性能,但由于对安全关键驾驶场景的暴露不足,确保在线RL策略的安全性仍具挑战性。现实交通状况的长尾特性使得通过常规采样难以遇到危险且罕见的交互场景,限制了RL策略学习鲁棒安全行为的能力。现有方法通过合成具有挑战性的场景或对抗性场景来提升训练多样性,但这些方法通常将场景生成目标与不断演化的策略分开优化,未明确建模生成的扰动与当前策略的弱点及学习需求之间的关联。本文提出用于在线强化学习安全自动驾驶的威胁引导、策略感知场景扰动(Threat-guided Policy-aware Scene Perturbation, TPSP)。TPSP引入策略感知场景编码器,以捕获策略行为与周围环境的交互,实现与当前策略对齐的场景扰动。基于该表示,TPSP选择性地扰动关键对象,而非对整个场景应用均匀修改。此外,我们开发了一种威胁引导优化策略,通过评估原始场景与扰动场景上策略 rollout 之间的威胁水平差异来评估扰动场景,引导生成具有更高训练价值的安全关键场景。综合实验表明,TPSP提升了安全学习效率,在NAVSIM v2数据集上利用约400万公里的模拟驾驶数据实现了出色的安全性能。消融研究验证了,策略感知的 targeted 扰动相比随机或无策略策略能提供更具信息性的安全关键体验,使策略在有限交互预算下实现更安全的驾驶。
英文摘要
Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes. The long-tailed nature of real-world traffic situations makes dangerous and rare interactions difficult to encounter through conventional sampling, limiting the ability of RL policies to learn robust safety behaviors. Existing methods improve training diversity by synthesizing challenging scenes or adversarial situations. However, these approaches typically optimize scene generation objectives separately from the evolving policy, without explicitly modeling how generated perturbations relate to the current policy's weaknesses and learning needs. In this paper, we propose Threat-guided Policy-aware Scene Perturbation (TPSP) for safe autonomous driving with online RL. TPSP introduces a policy-aware scene encoder to capture the interaction between policy behaviors and surrounding environments, enabling scene perturbation aligned with the current policy. Based on this representation, TPSP selectively perturbs critical objects rather than applying uniform modifications across the scene. Furthermore, we develop a threat-guided optimization strategy that evaluates perturbed scenes through threat-level differences between policy rollouts on original and perturbed scenes, guiding the generation of safety-critical scenes with higher training value. Comprehensive experiments demonstrate that TPSP improves safety learning efficiency, achieving strong safety performance on NAVSIM v2 with approximately 4 million kilometers of simulated driving data. Ablation studies verify that policy-aware targeted perturbations provide more informative safety-critical experiences than random or policy-unaware strategies, enabling safer driving under limited interaction budgets.
发表机构
- Yinwang Intelligent Technology Co., Ltd(银网智能科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。