arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21470cs.AI

风险感知占用:面向安全导向的端到端自动驾驶

Risk-Aware Occupancy for Safety-Oriented End-to-End Autonomous Driving

  • Tsinghua University(清华大学)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

Jiaxing Chen, Hengduo Zou, Yiren Zhao, Bolin Gao

AI总结:

本文提出风险感知占用表示及端到端网络ROIDrive,统一编码场景占用、交通约束和动态占用,以显式刻画规划风险,在nuScenes上分别减少52.9%和35.0%的开环碰撞。

AI中文摘要:

稀疏表示将端到端驾驶系统的环境感知表述为一组离散元素,如物体和车道线。这种表述在拥挤、遮挡的场景中处理非结构化障碍物、不确定区域和复杂交互时面临安全风险。在本文中,我们提出一种密集表示——风险感知占用,以显式且统一的方式刻画与规划相关的风险。它联合将全局场景占用、地图衍生的交通约束以及未来动态智能体占用编码到统一的鸟瞰图(BEV)中。该统一的BEV图在空间和时间维度上捕获轨迹规划的风险证据。我们设计了一个端到端网络ROIDrive来实现风险感知占用。它通过独立分支预测风险感知占用,并将其注入规划查询中以生成安全导向的轨迹。此外,为了量化安全问题,我们引入了基于nuScenes和occ3d-nuScenes构建的RiskOcc4D-nuScenes。我们的风险感知占用在nuScenes上,在UniAD指标下实现了52.9%的相对开环碰撞减少,在ST-P3指标下实现了35.0%的相对减少。

英文摘要:

Conventional end-to-end driving systems model the environment with sparse objects and lane elements. While efficient, this paradigm discards planning-critical information in crowded and occluded scenarios, particularly for unstructured obstacles, ambiguous free space, and complex interactions. We propose risk-aware occupancy, a dense BEV representation that explicitly fuses geometric occupancy, map-derived traffic constraints, and future dynamic-agent occupancy as complementary risk signals. Built upon this representation, we develop ROIDrive, an instance-centric end-to-end framework with a dedicated risk-aware occupancy branch. The predicted occupancy is tokenized via sliding-window sampling and injected into planning queries via cross-attention, while temporal query consistency mitigates unreliable flickering queries. We also contribute RiskOcc4D-nuScenes, a benchmark derived from nuScenes and Occ3D-nuScenes with four automated annotation pipelines for multi-dimensional risk supervision. Experiments on representative occupancy architectures verify the learnability and transferability of our representation. Integrated with GenAD, it reduces collision rates by 35.0% (UniAD metric) and 52.9% (ST-P3 metric), confirming the efficacy of the proposed representation modality.

补充信息

↑