UniJEPA:面向任务无关视觉世界建模的统一联合嵌入预测架构
UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling
浏览论文内容
中文总结 AI 辅助
UniJEPA是统一JEPAs,共享潜在空间联合学习图像级与视频级预测,抗坍塌,经后训练可零样本规划,在多基准上性能相当或更优,规划速度更快。
中文摘要 AI 辅助
联合嵌入预测架构(JEPAs)已成为在紧凑潜在空间中自监督学习世界模型的原则性框架,但现有方法存在碎片化问题:部分方法在潜在空间中预测单张图像的掩码部分(I-JEPA),其他方法学习预测全局光度变换(图像世界模型),而视频级JEPAs则预测未来时间状态,并经后训练用于动作条件规划(V-JEPA~2、DINO-World、DINO-WM)。这些目标被视为具有独立编码器、预测器和抗坍塌正则化器的不同方案,阻碍了单一模型统一图像级和视频级世界建模。我们提出UniJEPA,一种统一JEPAs,在共享潜在空间中联合学习光度预测(图像级变换)和时间预测(视频级下一状态动态)。由下一嵌入预测损失和高斯正则化器组成的单一端到端目标,产生可证明抗坍塌的编码器-预测器对,可从原始像素训练,无需EMA、停止梯度或预训练编码器。我们表明,相同潜在空间支持可控抽象:光度预测学习不变结构,而时间预测学习等变动态。在离线轨迹上进行动作条件后训练后,UniJEPA通过将目标特征视为预测目标实现零样本规划。在图像、视频和控制基准上,UniJEPA与任务特定JEPAs表现相当或更优,同时仅需单一损失超参数,且在相当精度下规划速度比生成式世界模型快数十倍。
英文摘要
Joint-Embedding Predictive Architectures (JEPAs) have emerged as a principled framework for self-supervised learning of world models in compact latent spaces, yet existing methods are fragmented: some predict masked parts of a single image in latent space (I-JEPA), others learn to predict global photometric transformations (Image World Models), while video-scale JEPAs predict future temporal states and are post-trained for action-conditioned planning (V-JEPA~2, DINO-World, DINO-WM). These objectives are treated as distinct recipes with separate encoders, predictors, and anti-collapse regularizers, hindering a single model from unifying image-level and video-level world modeling. We present UniJEPA, a unified JEPA that jointly learns photometric prediction (image-level transformations) and temporal prediction (video-level next-state dynamics) in one shared latent space. A single end-to-end objective, composed of a next-embedding prediction loss and a Gaussian regularizer, yields a provably anti-collapse encoder-predictor pair trainable from raw pixels without EMA, stop-gradient, or pre-trained encoders. We show that the same latent space supports controllable abstraction: photometric prediction learns invariant structure while temporal prediction learns equivariant dynamics. After action-conditioned post-training on offline trajectories, UniJEPA enables zero-shot planning by treating goal features as prediction targets. On image, video, and control benchmarks, UniJEPA matches or surpasses task-specific JEPAs while requiring a single loss hyperparameter, and plans up to tens of times faster than generative world models at comparable accuracy.