arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31383cs.ROcs.AIcs.CVcs.LG

引导端到端驾驶模型:端点约束轨迹优化

Guiding End-to-End Driving Models with Endpoint-Constrained Trajectory Optimization

  • University of Toronto(多伦多大学)
  • Vector Institute(向量研究所)
  • NVIDIA Research(英伟达研究院)
  • ELLIS Institute Tübingen(ELLIS 蒂宾根研究所)
  • KE:SAI

机构由 AI 辅助整理,请以论文原文为准。

Brayden Zhang, Mahsa Golchoubian, Igor Gilitschenski, Boris Ivanovic, Kashyap Chitta

AI总结:

针对端到端驾驶模型开环训练与闭环执行不匹配问题,提出端点约束优化(ECO)轻量级后处理层,在不改变预测端点前提下重塑中间路点,显著提升多种策略的闭环性能。

AI中文摘要:

端到端驾驶策略通常通过开环行为克隆进行训练,但部署在车辆上时最终必须在闭环中运行,这造成了训练与执行之间的根本性不匹配。除了通常研究的协变量偏移和因果混淆效应外,我们还识别出导致这种开环/闭环差距的一个补充因素:基于路点的监督和位移度量并不能确保中间轨迹在物理上连贯或易于控制器跟踪。我们观察到,这些不一致性主要集中在中间路点上,而预测的端点相对可靠。基于这一观察,我们引入了端点约束优化(ECO),这是一种轻量级后处理层,它将轨迹锚定到车辆已执行的轨迹历史,保留策略预测的端点,并重塑中间路点以提高可行性。ECO不需要地图、特权模拟器状态或额外训练,可以插入到广泛的路点生成策略及其控制器之间。在两个闭环模拟器中,它提高了所有六个评估的生成式和回归式驾驶策略的聚合闭环得分,且收益往往随着基础规划违反运动限制的频率增加而增加。在HUGSIM上,ECO将VaVAM从18.1提高到31.0 HD分数(+71%),在HUGSIM闭环驾驶挑战赛中取得第一名。同样,在AlpaSim上,ECO将VaVAM和DiffusionDrive的场景得分分别提高了123%和22%。这些结果表明,对于广泛的端到端驾驶模型,在不改变策略预测端点的情况下修复预测轨迹的中间几何结构,可以显著提高闭环性能。

英文摘要:

End-to-end driving policies are commonly trained through open-loop behavior cloning, yet they must ultimately operate in closed-loop when deployed on a vehicle, creating a fundamental mismatch between training and execution. Beyond the commonly studied effects of covariate shift and causal confusion, we identify a complementary factor for this open-loop/closed-loop gap: waypoint-based supervision and displacement metrics do not ensure that the intermediate trajectory is physically coherent or easy for the controller to track. We observe that these inconsistencies concentrate primarily at intermediate waypoints, while the predicted endpoint remains comparatively reliable. Based on this observation, we introduce Endpoint-Constrained Optimization (ECO), a lightweight postprocessing layer that anchors the trajectory to the vehicle's executed history, preserves the policy's predicted endpoint, and reshapes the intermediate waypoints to improve feasibility. ECO requires no map, privileged simulator state, or additional training, and can be inserted between a broad range of waypoint-emitting policies and their controllers. Across two closed-loop simulators, it improves the aggregate closed-loop score of all six evaluated generative and regression-based driving policies, and the gains tend to increase with how often the base plans violate motion limits. On HUGSIM, ECO improves VaVAM from 18.1 to 31.0 HD-Score (+71%), achieving 1st place on the HUGSIM Closed-Loop Driving Challenge. Similarly, on AlpaSim, ECO increases the scene scores of VaVAM and DiffusionDrive by 123% and 22%, respectively. These results show that for a broad collection of end-to-end driving models, repairing the intermediate geometry of predicted trajectories without changing the policy's predicted endpoint can substantially improve closed-loop performance.

↑