NavSafe-$\infty$:在逼真环境中对闭环驾驶安全性进行基准测试
NavSafe-$\infty$: Benchmarking Closed-Loop Driving Safety in Photorealistic Environments
浏览论文内容
中文总结 AI 辅助
本研究提出NavSafe-$\infty$闭环驾驶安全基准,含280个场景和28种事件类型,评估20种端到端策略,发现开环性能无法可靠迁移至闭环安全,并揭示常见补救措施的局限性,凸显开环基准的盲点。
中文摘要 AI 辅助
端到端(E2E)驾驶策略在开环(OL)基准测试上进展迅速,但开环评估无法揭示策略是否能承受复合误差、从失败中恢复,或与周围参与者安全交互。我们引入NavSafe-$\infty$,一个包含280个场景、涵盖28种事件类型的逼真闭环(CL)基准测试,每个场景在结构化交通安全分类法内定义了成功和失败标准,从而为交通事故、弱势道路使用者碰撞、交通违规和交通事件生成类别级能力评分。评估20种E2E策略后,我们发现开环收益并不能可靠地转化为闭环安全性。进一步分析两种常见补救措施表明,被动演示扰动主要仅在闭环轨迹接近其扰动训练状态时有所帮助,而开环强化学习微调表现出奖励黑客行为,通过牺牲安全边际换取自我进展,闭环反馈将这种权衡放大为复合的安全关键错误。这些结果共同揭示了开环基准在指示闭环安全成功方面的盲点。该基准以及用于可定制事件策划和策略诊断的可扩展工具箱将开源并持续维护,以促进未来研究。
英文摘要
End-to-end (E2E) driving policies have advanced rapidly on open-loop (OL) benchmarks, yet OL evaluation cannot reveal whether a policy can withstand compounding errors, recover from failures, or interact safely with surrounding actors. We introduce NavSafe-$\infty$, a photorealistic closed-loop (CL) benchmark comprising 280 scenarios spanning 28 event types, each with success and failure criteria defined within a structured traffic-safety taxonomy, yielding category-level capability scores for Traffic Crashes, Vulnerable Road User Crashes, Traffic Violations, and Traffic Incidents. After evaluating 20 E2E policies, we find that OL gains do not reliably transfer to CL safety. Analysis of two common remedies reveals that (1) passive demonstration perturbation helps mainly when CL rollouts stay near their perturbed training states, and (2) OL reinforcement-learning fine-tuning exhibits reward hacking by trading safety margin for ego progress, which CL feedback amplifies into compounding safety-critical errors. Together, these results demonstrate the blind spot of OL benchmarks indicating CL safety success. The benchmark and an extensible toolbox for customizable event curation and policy diagnosis will be open-sourced to facilitate future research.
发表机构
- UCLA(加州大学洛杉矶分校)
- UCSD(加州大学圣迭戈分校)
- Toyota Research Institute(丰田研究所)
机构由 AI 辅助整理,请以论文原文为准。