arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30436cs.ROcs.CV

WALT:学习面向自动驾驶的世界模型对齐潜在轨迹

WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving

  • The Hong Kong University of Science and Technology(香港科技大学)
  • Horizon Robotics(地平线机器人)
  • The Chinese University of Hong Kong(香港中文大学)
  • Nanjing University of Posts and Telecommunications(南京邮电大学)
  • Nankai University(南开大学)

机构由 AI 辅助整理,请以论文原文为准。

Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin

AI总结:

针对视觉预测与轨迹规划不匹配问题,提出WALT,通过双分支自编码器将路径点映射为潜在轨迹并迁移冻结世界模型语义知识,在NAVSIM上提升PDMS/EPDMS并降低30.5%计算量。

AI中文摘要:

驾驶世界模型从视觉观测中学习周围环境的丰富预测性表征,然而准确的视觉预测并不一定能转化为有效的轨迹规划。我们认为,关键瓶颈在于视觉世界状态与原始几何轨迹之间的不匹配,这可能限制规划器利用世界模型编码的与动作相关的语义信息。为解决这一问题,我们提出了世界模型对齐潜在轨迹(World-Model Alignment for Latent Trajectories, WALT),该方法通过从冻结的预训练驾驶世界模型中迁移信息来学习一个紧凑的生成式轨迹潜在空间,而无需修改世界模型本身。WALT并非直接生成原始路径点,而是通过一个双分支轨迹自编码器将其映射为紧凑表征,并将冻结的视觉世界模型中的语义知识迁移到该轨迹空间中,从而鼓励所学动作表征捕获与未来运动和规划相关的场景级线索。除所提出的公式外,我们还基于联合嵌入预测架构(Joint-Embedding Predictive Architectures, JEPA)和表征对齐(Representation Alignment, REPA)后的特征对齐系统研究了潜在学习,以探究仅基于轨迹的表征学习如何影响下游规划。我们在NAVSIM基准上评估了WALT。相对于原始路径点基线,WALT在NAVSIMv1上将PDMS从89.4提升至89.8,在NAVSIMv2上将EPDMS从87.3提升至87.9,同时将轨迹规划器的FLOPs降低了30.5%。这些结果表明,在提取与动作相关信息的同时保留世界表征,为基于世界模型的轨迹规划提供了一种有效接口。

英文摘要:

Driving world models learn rich predictive representations of the surrounding environment from visual observations, yet accurate visual prediction does not necessarily translate into effective trajectory planning. We argue that a key bottleneck lies in the mismatch between visual world states and raw geometric trajectories, which may limit the planner's ability to exploit action-relevant semantics encoded by the world model. To address this issue, we propose World-Model Alignment for Latent Trajectories (WALT), which learns a compact generative trajectory latent space by transferring information from a frozen pretrained driving world model without modifying the world model itself. Rather than directly generating raw waypoints, WALT maps them into compact representations through a dual-branch trajectory autoencoder and transfers semantic knowledge from the frozen visual world model into this trajectory space, encouraging the learned action representation to capture scene-level cues relevant to future motion and planning. Beyond our proposed formulation, we systematically study latent learning based on Joint-Embedding Predictive Architectures (JEPA) and feature alignment following Representation Alignment (REPA) to investigate how trajectory-only representation learning affects downstream planning. We evaluate WALT on the NAVSIM benchmarks. Relative to the raw-waypoint baseline, WALT improves PDMS from 89.4 to 89.8 on NAVSIMv1 and EPDMS from 87.3 to 87.9 on NAVSIMv2 while reducing trajectory planner FLOPs by 30.5%. These results suggest that preserving world representations while extracting action-relevant information provides an effective interface for world-model-based trajectory planning.

相关深度报道

↑