Brain-SAD:一种具有动态恐惧导向约束的双策略脑启发式安全自动驾驶控制框架
Brain-SAD: A Brain-Inspired Safe Autonomous Driving Control Framework with Dynamic Fear-Oriented Constraint on Dual-Policy
查看机构详情
- Nanjing University of Information Science & Technology(南京信息工程大学)
- University of Electronic Science and Technology of China(电子科技大学)
- Jiangsu Ocean University(江苏海洋大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对现有约束强化学习在安全自动驾驶中约束静态化的问题,提出脑启发式Brain-SAD框架,通过动态恐惧信号生成双策略及动态约束,提升成功率与可靠性。
中文摘要 AI 辅助
约束强化学习近年来在安全自动驾驶领域受到越来越多的关注,其一般机制是在保持整体动作风险有界的同时最大化期望回报。通过这种方式,自动驾驶中出现的安全问题可以通过约束动作来缓解。然而,现有的约束强化学习方法在施加的约束上仍然缺乏动态性。例如,现有的原始-对偶/软约束方法所采用的动作成本通常被定义为静态的状态到成本映射,而硬约束方法中的安全动作投影依赖于从离线演示中估计的固定可行区域边界的静态投影。上述缺点将施加的约束与训练场景紧密耦合,使得自动驾驶策略难以处理不同的交互场景,原因是状态级动作成本不当以及静态投影边界。因此,在本文中,我们提出了Brain-SAD,一种具有动态恐惧导向约束的脑启发式安全自动驾驶控制框架。通过感知当前车辆交互场景,Brain-SAD生成动态恐惧信号作为恐惧反应,以在线决定用于常规交互的长期策略或用于紧急碰撞防御的短期策略。在这两种策略中,上述恐惧反应将被构建为动态恐惧约束,分别反映与动作影响直接耦合的整体恐惧成本,以及源自不同风险邻居的可行区域动态恐惧边界,这两者都将反过来用于在线策略优化。实验结果表明,Brain-SAD优于现有方法,在更短的任务完成和碰撞恢复时间内实现了更高的成功率,并在复杂度波动的连续交叉口中表现出更强的可靠性。
英文摘要
Constrained Reinforcement Learning has recently gained increasing attention in the field of Safe Autonomous Driving, where the general mechanism is to maximize the expected reward while keeping the overall action risk bounded. In this way, the safety issues arising in AD can be mitigated through constrained actions. However, existing Constrained RL methods still lack dynamics on the imposed constraints. For instance, the action cost adopted by the existing Primal-Dual/soft-constrained methods is often defined as static state-to-cost mapping, and the safe-action projection in hard-constrained methods relies on the static projection with the fixed feasible region boundary estimated from offline demonstrations. The above drawback tightly couples the imposed constraints to the training scenarios, leaving the AD policy hard to handle different interaction scenarios, due to the improper state-level action-cost and the static projection boundary. Consequently, in this paper, we propose Brain-SAD, a brain-inspired safe autonomous driving control framework with dynamic fear-oriented constraints. By perceiving the current vehicle-interaction scene, Brain-SAD generates dynamic fear signal as fear reaction to online decide long-term policy for regular interaction or short-term policy for urgent-collision defense. In such two policy, the above fear-reaction will be constructed as the dynamic fear constraints, respectively reflecting the overall fear cost directly coupled with action-impact, and the dynamic fear boundary of the feasible region derived from different risky neighbors, both of which will in turn serve for the online policy optimization. Experimental results show that Brain-SAD outperforms existing methods, achieving higher success rate in shorter task-completion and collision-recovery time, and exhibits stronger reliability across continuous intersections of fluctuating complexity.