发表机构
Hanyang University(汉阳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PAVER通过稀疏动作条件目标进行规划对齐的BEV编码器预训练,无需任务标注或密集重建,显著降低碰撞率并提升驾驶性能。
AI 中文摘要
端到端驾驶需要与规划相关的鸟瞰图(BEV)表征,但现有的预训练方法通常依赖任务标注或密集场景重建。我们提出了PAVER,即规划对齐的BEV编码器预训练方法。从单次LiDAR扫描中,PAVER构建稀疏的风险和未知目标,描述沿基于规则的自我运动轨迹上的占据和未观测证据。一个10K参数的预测头从基于动作状态条件的掩码BEV特征中预测这些目标,将监督引导至候选运动的几何约束上。预训练不需要驾驶任务标注或密集重建。仅迁移BEV编码器,保留下游架构和仅相机推理。在nuScenes上,PAVER将VAD-Tiny的平均碰撞率从0.51%降低至0.19%,同时改善了规划L2、运动预测、检测和建图。所选的VAD-Tiny和VAD-Base训练计划相比从头训练(包括预训练)估计节省约36%的总训练时间。在Bench2Drive Town05 Long上,PAVER将UniAD-Tiny的闭环驾驶得分从48.45提升至58.79。项目页面可在该https URL获取。
英文摘要
End-to-end driving requires planning-relevant bird's-eye-view (BEV) representations, but existing pretraining approaches often rely on task annotations or dense scene reconstruction. We introduce PAVER, Planning-Aligned BEV Encoder Pretraining. From a single LiDAR sweep, PAVER constructs sparse risk and unknown targets describing occupied and unobserved evidence along rule-based ego motions. A 10K-parameter head predicts these targets from masked BEV features conditioned on the action state, directing supervision toward geometric constraints on candidate motions. Pretraining requires no driving-task annotations or dense reconstruction. Only the BEV encoder is transferred, preserving the downstream architecture and camera-only inference. On nuScenes, PAVER reduces VAD-Tiny's average collision rate from 0.51% to 0.19%, while improving planning L2, motion prediction, detection, and mapping. The selected VAD-Tiny and VAD-Base schedules use about 36% less estimated total training time than scratch training, including pretraining. On Bench2Drive Town05 Long, PAVER improves UniAD-Tiny's closed-loop Driving Score from 48.45 to 58.79. The project page is available at https://archiiive99.github.io/PAVER.