JEPA Policy:通过配对动作与未来表示预测实现无扩散模仿学习
JEPA Policy: Diffusion-Free Imitation Learning via Paired Action and Future Representation Prediction
浏览论文内容
中文总结 AI 辅助
JEPA Policy通过配对动作与未来表示预测,在共享Transformer中交互,以无扩散方式提升模仿学习性能,降低延迟并优于现有基线。
中文摘要 AI 辅助
标准行为克隆在监督动作时,并未显式约束与每个演示动作块配对的未来表示。我们提出JEPA Policy,一种无扩散框架,将动作块及其观测到的未来表示作为配对训练目标。动作与未来表示标记在共享Transformer中交互,并通过两次前向传播进行细化。因此,未来预测能够塑造用于生成动作的表示。双分支和梯度路由控制将增益归因于这种共享拓扑结构,而非仅归因于辅助预测头。在九个模拟任务中,JEPA Policy相比仅动作的MIP基线提高了平均成功率,并在评估配置下优于Diffusion Policy,同时仅增加MIP模型延迟0.29毫秒。一项包含五个任务、630次试验的物理机器人研究得出了相同的汇总排名。进一步审计发现,在动作监督下没有出现完全的表示崩溃,并在未来预测误差中识别出任务条件化的失败排名信号。这些结果支持将配对未来表示监督作为一种无需迭代生成采样的低延迟视觉运动模仿的实用方法。
英文摘要
Standard behavior cloning supervises actions without explicitly constraining the future representation paired with each demonstrated action chunk. We introduce JEPA Policy, a diffusion-free framework that uses the action chunk and its observed future representation as paired training targets. Action and future-representation tokens interact in a shared Transformer and are refined through two forward passes. Future prediction can therefore shape the representation used to generate actions. Dual-branch and gradient-routing controls attribute the gain to this shared topology rather than to an auxiliary prediction head alone. Across nine simulated tasks, JEPA Policy improves mean success over the action-only MIP baseline and outperforms Diffusion Policy under the evaluated configurations, while adding 0.29 ms to MIP's model latency. A five-task, 630-episode physical-robot study produces the same pooled ranking. Further audits find no complete representation collapse under action supervision and identify a task-conditioned failure-ranking signal in future-prediction error. These results support paired future-representation supervision as a practical approach to low-latency visuomotor imitation without iterative generative sampling.
发表机构
- Anyverse Dynamics
机构由 AI 辅助整理,请以论文原文为准。