TwinJEPA:面向目标条件控制的动作偏好预测表示
TwinJEPA: Action-Preferred Predictive Representations for Goal-Conditioned Control
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
TwinJEPA通过离线挖掘动作偏好监督增强JEPA表示学习,在长时程控制中同时保留时间结构与局部动作区分,实现离线零样本控制性能提升。
AI中文摘要:
联合嵌入预测架构(JEPAs)是一种通过预测未来潜在状态而不重构观测来学习表示的有前景范式。近期工作已将JEPA风格的潜在预测适配到离线零样本控制中,通过动作条件的时间预测,使表示能够从固定的行为数据中捕捉长时程动态。然而,转移级别的预测目标孤立地监督动作,并且在相似状态允许不同目标条件结果时,提供关于哪些动作更可取的有限信息。我们引入TwinJEPA,一个通过用离线挖掘的动作偏好监督增强基于JEPA的控制来学习动作偏好预测表示的框架。TwinJEPA在离线轨迹中识别近似匹配的状态,并通过目标条件奖励重标记构建偏好对。然后它学习两个互补目标:奖励差距回归,保留结果差异的幅度;以及偏好分类,捕捉替代动作的相对排序。这两个目标仅在训练期间使用,不产生额外的推理时间成本。我们在具有状态和像素观测的长时程导航和连续控制基准上评估TwinJEPA。TwinJEPA在所有匹配的状态基评估中产生正的基准级平均差异,而跨领域和观测模态的分析表明,当局部动作替代提供更具信息性的结果对比时,往往会出现更大的收益。结果表明,局部动作比较监督可以通过鼓励基于JEPA的表示同时保留长时程时间结构和局部观测动作间的细粒度区分,来补充时间预测,从而用于离线零样本控制。
英文摘要:
Joint-Embedding Predictive Architectures (JEPAs) are a promising paradigm for representation learning by predicting future latent states without reconstructing observations. Recent work has adapted JEPA-style latent prediction to offline zero-shot control through action-conditioned temporal prediction, enabling representations to capture long-horizon dynamics from fixed behavioral data. However, transition-level predictive objectives supervise actions in isolation and provide limited information about which actions are preferable when similar states admit different goal-conditioned outcomes. We introduce TwinJEPA, a framework for learning action-preferred predictive representations by augmenting JEPA-based control with offline-mined action-preference supervision. TwinJEPA identifies approximately matched states across offline trajectories and constructs preference pairs via goal-conditioned reward relabeling. It then learns two complementary objectives: reward-gap regression, which preserves the magnitude of outcome differences, and preference classification, which captures the relative ordering of alternative actions. Both objectives are used only during training and incur no additional inference-time cost. We evaluate TwinJEPA on long-horizon navigation and continuous-control benchmarks with state- and pixel-based observations. TwinJEPA yields positive benchmark-level mean differences across all matched state-based evaluations, while analyses across domains and observation modalities indicate that larger gains tend to arise when local action alternatives provide more informative outcome contrasts. The results suggest that local action-comparison supervision can complement temporal prediction by encouraging JEPA-based representations to retain both long-horizon temporal structure and fine-grained distinctions among locally observed actions for offline zero-shot control.