发表机构
Hong Kong Baptist University(香港浸会大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出时间距离联合嵌入预测架构(TD-JEPA),保留LeWM编码器-预测器主干,从无奖励轨迹挖掘有向时间成本。该方法缩小JEPA世界模型规划器训练-规划差距,在多个环境中表现优于LeWM及其他基线,通过挖掘成本提升成功率和分数。
AI 中文摘要
联合嵌入预测架构(JEPAs)通过在表示空间中进行预测来学习世界模型,而非重建像素,是潜在模型预测控制的自然基础。JEPA式训练优化短期潜在预测,而规划需要按目标进展对想象的未来进行多步排序。先前的JEPA规划器常从嵌入几何继承排序,通常是潜在欧几里得距离,这是表示学习的副产品而非从日志中挖掘的进展成本。我们提出时间距离JEPA(TD-JEPA),保留LeWM编码器-预测器主干,并从无奖励轨迹中挖掘有向时间成本:同轨迹步序提供正目标,交叉轨迹对作为启发式负目标,展开一致性项匹配规划器视野。挖掘的监督有两个作用:当进展是拓扑结构时作为部署的规划成本,当接触几何占主导时作为改善欧几里得规划的表示信号。在锁定评估下,部署挖掘的成本将双房间成功率提高到100.0%,而LeWM为97.4%,在相同时间训练的检查点上共享欧几里得规划使OGB-Cube比LeWM提高14.2分并改善Push-T。与LeWM及并发RC-aux基线相比,TD-JEPA在每个环境中都匹配或超过这两种方法。消融实验表明有向头、交叉轨迹负目标和展开一致性都有贡献。TD-JEPA通过在离线日志中发现时间进展结构并与规划时部署共同设计成本形式,缩小了JEPA世界模型规划器的训练-规划差距。代码可在该https URL获取。
英文摘要
Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rather than reconstructing pixels, making them a natural backbone for latent model predictive control from offline demonstration logs. JEPA-style training optimizes short-horizon latent prediction, whereas planning requires a multi-step ranking of imagined futures by goal progress. Prior JEPA planners often inherit that ranking from embedding geometry, typically latent Euclidean distance, which arises as a byproduct of representation learning rather than as a progress cost mined from the logs. We propose Temporal-Distance-JEPA, which retains the LeWM encoder--predictor backbone and mines a directed temporal cost from reward-free trajectories: same-trajectory step order supplies positive targets, cross-trajectory pairs act as heuristic negatives, and a rollout-consistency term matches the planner horizon. The mined supervision serves two roles: as the deployed planning cost when progress is topological, and as a representation signal that improves Euclidean planning when contact geometry dominates. Under locked evaluation, deploying the mined cost raises Two-Room success to 100.0% versus LeWM's 97.4%, while shared Euclidean planning on the same temporally trained checkpoint raises OGB-Cube by 14.2 points over LeWM and improves Push-T. Against LeWM and the concurrent RC-aux baseline under locked evaluation, Temporal-Distance-JEPA matches or exceeds both methods on every environment. Ablations show that the directed head, cross-trajectory negatives, and rollout consistency each contribute. Temporal-Distance-JEPA narrows the train--plan gap for JEPA world-model planners by discovering temporal progress structure in offline logs and co-designing cost form with plan-time deployment. Code is available at https://github.com/HKBU-KnowComp/Temporal-Distance-JEPA.