发表机构
University of Toronto; Zhejiang University; Tencent Jarvis Lab; Mila & Université de Montréal; Samsung SAIL; Tsinghua University(多伦多大学; 浙江大学; 腾讯Jarvis实验室; Mila与蒙特利尔大学; 三星SAIL; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FlexiWorld提出基于JEPA的世界模型,结合混合跨度监督与可变长度动作块,配合ARCEM规划方法,在四个基准上实现89.29%的平均成功率,优于基线,并支持灵活块长度加速规划。
AI 中文摘要
潜在世界模型通过跨越多个原始步骤的动作块预测未来状态,以支持目标导向的规划。现有方法通常使用固定长度的动作块,并且要么省略目标条件动作生成,要么将监督限制在短目标跨度内。我们提出FlexiWorld,一种基于JEPA的世界模型,结合混合跨度目标监督与可变长度动作块,以改善长时程控制。在训练期间,我们采样不同的目标跨度,并随机将动作划分为可变长度的块。我们联合训练世界模型与一个因果动作编码器(该编码器嵌入可变长度块)以及一个自回归动作生成器(该生成器顺序生成原始动作)。学生强制通过训练生成的动作前缀来减少暴露偏差。对于规划,Actor-Residual Cross-Entropy Method (ARCEM) 结合了动作残差搜索、块内自回归反馈和块边界潜在预测。在四个基准和不同目标距离上,FlexiWorld与ARCEM实现了89.29%的平均成功率,而最强基线为83.98%。PushT消融实验表明,混合跨度监督、可变长度块和学生强制改善了直接控制。无需重新训练,FlexiWorld支持不同的规划块长度:更长的块平均加速ARCEM约1.3倍,同时保持相当的平均成功率。
英文摘要
Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or limit their supervision to short goal spans. We introduce FlexiWorld, a JEPA-based world model that combines mixed-span goal supervision with variable-length action chunks to improve long-horizon control. During training, we sample varying goal spans and randomly partition the actions into variable-length chunks. We jointly train the world model with a causal action encoder that embeds variable-length chunks and an autoregressive actor that generates primitive actions sequentially. Student Forcing reduces exposure bias by training on generated action prefixes. For planning, Actor-Residual Cross-Entropy Method (ARCEM) combines action-residual search with within-chunk autoregressive feedback and chunk-boundary latent prediction. Across four benchmarks and goal distances, FlexiWorld with ARCEM achieves 89.29% mean success, compared with 83.98% for the strongest baseline. PushT ablations show improved direct control from mixed-span supervision, variable-length chunks, and Student Forcing. Without retraining, FlexiWorld supports different planning chunk lengths: longer chunks accelerate ARCEM by approximately $1.3\times$ on average while maintaining comparable average success.
Comments25 pages, 12 figures. Project page: https://shidu-ren.github.io/FlexiWorld-Project-Page/