发表机构
Nanyang Technological University; The University of Queensland(南洋理工大学; 昆士兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对JEPA世界模型确定性自回归预测器易累积误差且对扰动敏感的问题,提出Flow-JEPA,采用条件流匹配实现随机轨迹级预测,显著提升了干净及含噪观测下的平均成功率。
AI 中文摘要
联合嵌入预测架构(JEPAs)在学习紧凑预测表示方面展现出强大潜力,LeWorldModel(LeWM)将该范式扩展为从像素出发的无重建潜在世界建模。然而,其确定性自回归预测器通过重复的单步转移生成未来状态,会累积误差且对任务无关的视觉扰动敏感。本研究提出Flow-JEPA(F-JEPA),一种条件流匹配动力学模型,该模型基于当前观测与动作联合生成未来潜在状态序列。高斯分布作为流源,在学习将受扰动潜在轨迹迁移至干净未来表示时,向向量场暴露这些轨迹。该方案保留无重建JEPA框架,同时将逐点转移回归替换为随机轨迹级预测。F-JEPA在干净观测下将平均成功率从86%提升至92%,在含噪条件下从67%提升至86%,表明条件流匹配为JEPA世界模型中的确定性自回归动力学提供了有前景的替代方案。
英文摘要
Joint-Embedding Predictive Architectures (JEPAs) provide a powerful framework for latent world modeling and planning in a reconstruction-free manner. Although numerous JEPA-based approaches have been proposed to mitigate representation collapse, our experiments on localized, out-of-distribution visual noise reveal that performance degradation remains pronounced and unresolved. We propose Flow-JEPA (F-JEPA), a flow-based latent dynamics model that jointly generates a sequence of future latent states conditioned on the current observation and actions. A Gaussian distribution serves as the flow source, exposing the vector field to perturbed latent trajectories as it learns to transport them toward clean future representations. This formulation retains the reconstruction-free JEPA framework while switching from pointwise transition regression to stochastic trajectory-level prediction. F-JEPA raises mean success from $86\%$ to $92\%$ under clean observations and from $67\%$ to $86\%$ under noisy conditions. Further evaluations over varying perturbation severity and inference settings show that the performance advantage persists across a broad range of conditions. These results suggest that conditional flow matching provides a promising alternative to deterministic autoregressive prediction as a dynamics formulation in JEPA world models.