发表机构
School of Vehicle and Mobility, Tsinghua University; State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University(清华大学车辆与运载学院; 清华大学智能网联汽车与交通国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对端到端轨迹规划器问题,提出DRIFT固定深度规划器,结合轨迹潜在空间漂移与提议聚合,基于视觉编码器特征生成提议,经聚合头预测最终轨迹,在导航测试中有良好表现,证明该方法为多假设运动规划提供高效设计。
AI 中文摘要
端到端轨迹规划器需在实时约束下表示多种合理驾驶行为并生成单一可执行轨迹。基于提议的方法通过生成多个候选来解决模糊性,但将提议集转换为最终计划仍是关键设计问题。我们提出DRIFT,一种固定深度规划器,它在紧凑轨迹潜在空间中结合一步漂移与场景感知提议聚合。基于预训练视觉编码器的特征,DRIFT解码器在单次批处理中生成48个提议特征,在α = 0.5时有32个样本,α = 0.9时有16个样本。轻量级聚合头将这些特征与场景、导航和自我状态信息集成,直接预测最终轨迹,无需轨迹级质量标签进行聚合。其输出通过专家轨迹模仿和地图衍生边界正则化器训练,该正则化器惩罚可驾驶多边形外和其边界附近内部的航点。在NAVSIM导航测试中,DRIFT实现了89.6 PDMS和90.4 EPDMS,在可比方法中具有很强的可驾驶区域合规性和自我进展。提议生成和聚合模块在NVIDIA RTX 4090上运行10.82毫秒,包括视觉主干的全模型推理需要66.43毫秒。这些结果表明一步潜在提议生成和直接聚合为多假设运动规划提供了高效设计。
英文摘要
End-to-end trajectory planners need to represent multiple plausible driving behaviors while producing a single executable trajectory under real-time constraints. Proposal-based approaches address this ambiguity by generating multiple candidates, but converting the proposal set into a final plan remains a key design problem. We present DRIFT, a fixed-depth planner that combines one-step drifting in a compact trajectory latent space with scene-aware proposal aggregation. Conditioned on features from a pretrained visual encoder, the DRIFT Decoder generates 48 proposal features in a single batched pass, with 32 samples at alpha=0.5 and 16 samples at alpha=0.9. A lightweight Aggregation Head integrates these features with scene, navigation, and ego-state information and directly predicts the final trajectory without requiring trajectory-level quality labels for aggregation. Its output is trained with expert-trajectory imitation and a map-derived boundary regularizer that penalizes waypoints outside the drivable polygon and inside waypoints near its boundary. On NAVSIM navtest, DRIFT achieves 89.6 PDMS and 90.4 EPDMS, with strong drivable-area compliance and ego progress among the methods compared. The proposal-generation and aggregation module runs in 10.82 ms on an NVIDIA RTX 4090, while full-model inference including the visual backbone takes 66.43 ms. These results show that one-step latent proposal generation and direct aggregation provide an efficient design for multi-hypothesis motion planning.
Comments8 pages, 3 figures, 4 tables. Under review at IEEE RAL