arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14125cs.AI

Traj-LeWM:通过潜在轨迹代价实现路径感知的世界模型规划

Traj-LeWM: Path-Aware World-Model Planning via Latent Trajectory Cost

  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • Shanghai Jiao Tong University(上海交通大学)
  • Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)
  • The University of Hong Kong(香港大学)
  • INFIFORCE
  • University of Science and Technology of China(中国科学技术大学)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiaodi Huang, Ziyi Ding, Jingtian Wan, Yuchen Liu, Yuan Zhang, Xiao-Ping Zhang, Jiayu Chen, Zhang Zhang, Tao Huang

AI总结:

本文针对LeWM的局限提出Traj-LeWM,通过引入潜在轨迹代价结合终点距离的联合评分,在多个机器人与导航任务上实现性能提升,验证了轨迹级信息的互补作用。

AI中文摘要:

LeWM是一种轻量级视觉世界模型,可从像素中端到端学习潜在动力学,并通过候选动作序列的预测终点与目标的距离对其进行排序。但LeWM存在两个局限:一是训练时仅学习局部下一步转移,未评估相对于任务目标的完整轨迹;二是规划时仅通过预测终点距离排序候选,而模型预测可能与实际执行结果存在差异,预测终点最接近目标的候选在环境中执行时未必表现最佳,完整预测轨迹的演化可提供超出终点距离的补充信息。为解决这些局限,本文提出Traj-LeWM,它保留LeWM的局部动力学目标和终点评分,同时引入目标条件潜在轨迹代价(LTC),将轨迹级信息聚合为补充信号。训练时,基于LTC的轨迹偏好监督与下一步预测共同塑造共享表征;规划时,LTC与终点距离结合,将中间路径信息纳入候选排序。采用终点加LTC的联合评分,Traj-LeWM在Push-T、OGBench-Cube、Reacher和Two-Room上分别比LeWM高出3、14、7、7个百分点,控制实验和消融研究进一步验证了轨迹级表征塑造和路径感知候选排序的互补作用。

英文摘要:

LeWM is a lightweight visual world model that learns latent dynamics end-to-end from pixels and ranks candidate action sequences by the distance between their predicted endpoints and the goal. However, LeWM has two limitations. First, during training, it learns local next-step transitions without evaluating complete trajectories relative to the task goal. Second, during planning, it ranks candidates solely by predicted endpoint distance. Because model predictions may differ from actual execution outcomes, the candidate whose predicted endpoint is closest to the goal may not perform best when executed in the environment. The evolution of the complete predicted trajectory can therefore provide complementary information beyond endpoint distance. To address these limitations, we propose Traj-LeWM, which retains LeWM's local-dynamics objective and endpoint score while introducing a goal-conditioned latent trajectory cost (LTC) that aggregates trajectory-level information as a complementary signal. During training, LTC-based trajectory-preference supervision complements next-step prediction in shaping the shared representation. During planning, LTC is combined with endpoint distance to incorporate intermediate-path information into candidate ranking. With joint endpoint-plus-LTC scoring, Traj-LeWM outperforms LeWM on Push-T, OGBench-Cube, Reacher, and Two-Room by $3$, $14$, $7$, and $7$ percentage points, respectively. Controlled experiments and ablations further verify the complementary roles of trajectory-level representation shaping and path-aware candidate ranking.

↑