学习在规划中解释:面向自动驾驶的规则对齐扩散规划
Learning to Explain While Planning: Rule-Aligned Diffusion Planning for Autonomous Driving
浏览论文内容
中文总结 AI 辅助
针对扩散规划器缺乏显式规则建模和解释的问题,提出规则对齐扩散规划器(RADP)将可微驾驶规则融入训练,并引入规则压力归因(RPA)在线估计规则作用,实验验证了安全关键场景下的闭环性能提升和规则级可解释性。
中文摘要 AI 辅助
扩散规划器在生成多模态轨迹方面表现出强大的能力。然而,现有方法主要依赖专家演示来拟合轨迹分布,学习场景、行为和轨迹之间的统计相关性,而没有显式建模驾驶规则。在专家数据稀缺的长尾场景中,缺乏可模仿的行为可能导致轨迹违反安全或合规要求。此外,其生成过程缺乏规则级解释,难以确定哪些规则驱动轨迹调整、何时生效以及作用强度,从而限制了故障诊断、安全验证和针对性改进。为解决这些局限,我们提出了规则对齐扩散规划器(RADP),在训练期间将可微驾驶规则纳入扩散目标,将规则知识转化为超越有限演示的内在行为原则。我们进一步引入了规则压力归因(RPA),该方法从预测轨迹的规则损失梯度构建监督信号,并采用轻量级归因头在线估计每条规则施加的优化压力。为评估这些归因的闭环行为相关性,我们提出了一种时间风险对齐协议,评估当前规则压力是否反映后续闭环执行中的相应风险。在nuPlan上的实验表明,RADP在具有挑战性的安全关键场景中改善了闭环规划,而RPA与后续规则特定风险表现出一致的时间对齐,验证了内在规则学习和规则级可解释性。
英文摘要
Diffusion planners exhibit strong capabilities in generating multimodal trajectories. However, existing methods primarily rely on expert demonstrations to fit trajectory distributions, learning statistical correlations among scenes, behaviors, and trajectories without explicitly modeling driving rules. In long-tail scenarios where expert data are scarce, the lack of behaviors to imitate may lead to trajectories that violate safety or compliance requirements. Moreover, their generation process lacks rule-level explanations, making it difficult to determine which rules drive trajectory adjustments, when they take effect, and how strongly they act, thereby limiting failure diagnosis, safety validation, and targeted improvement. To address these limitations, we propose the Rule-Aligned Diffusion Planner (RADP), which incorporates differentiable driving rules into the diffusion objective during training, turning rule knowledge into intrinsic behavioral principles beyond finite demonstrations. We further introduce Rule-Pressure Attribution (RPA), which constructs supervision signals from gradients of rule losses with respect to predicted trajectories and employs a lightweight attribution head to estimate the optimization pressure exerted by each rule online. To assess the closed-loop behavioral relevance of these attributions, we propose a temporal risk-alignment protocol that evaluates whether current rule pressures reflect corresponding risks during subsequent closed-loop execution. Experiments on nuPlan show that RADP improves closed-loop planning in challenging safety-critical scenarios, while RPA exhibits consistent temporal alignment with subsequent rule-specific risks, validating both intrinsic rule learning and rule-level interpretability.
发表机构
- Beijing Institute of Technology(北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。