arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Do-JEPA:从掩蔽到潜在世界模型中的干预

Do-JEPA: From Masking to Intervention in Latent World Models

Hossein Resani, Javen Qinfeng Shi

arXiv 2609.37378首次发表:更新:

发表机构

Australian Institute for Machine Learning, Adelaide University; Responsible AI Research Centre(阿德莱德大学澳大利亚机器学习研究所; 负责任人工智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Do-JEPA通过干预物理世界而非模型所见内容,训练潜在世界模型预测动作因果效应,在多个基准上显著降低效应误差并提升因果发现能力。

AI 中文摘要

潜在世界模型被训练用于预测接下来会发生什么,因此其目标中没有任何内容能够区分某个动作所导致的结果与仅仅与之同时发生的结果。诸如C-JEPA之类的对象掩蔽模型对预测器所能看到的内容进行干预;而我们则对物理上实际发生的事情进行干预。从保存的单个模拟器状态出发,我们在动作$a$和参考动作$a_{\varnothing}$下运行动力学,并训练模型预测两种潜在未来之间的差异$\Delta z=z^{a}-z^{a_{\varnothing}}$。由此产生的目标Do-JEPA包含一个效应损失、一个支持损失(动作进入之处)、一个传播损失(其效应传播之处)以及不变性损失(必须保持不变的内容)。在一个具有对象对齐变量的合成系统中,支持监督在99.95%的测试案例中找到了被直接干预的对象,而稀疏动作掩蔽则在每个案例中都将动作发送到一个干扰槽中,响应起始监督恢复了环形传播图(边缘AUROC 0.975对0.624)。从像素出发,效应损失优于在完全相同数据上训练的对照组:在端到端LeWM模型上将潜在效应误差降低了28.4%,在自然动作序列上训练和测试时将物理效应误差降低了13.5%,并且在三个独立生成的CausalWorld基准上,在物理变化下将响应效应误差降低了约20%,并将预测效应的潜在上下文敏感性降低了66%。从头开始训练会牺牲事实准确性;用其微调现有模型则消除了这一代价。总之,这些结果表明,对世界进行干预,而非对模型所见内容进行干预,有助于潜在世界模型预测其动作所导致的结果。

英文摘要

Latent world models are trained to predict what happens next, so nothing in their objective separates what an action caused from what merely co-occurred with it. Object-masking models such as C-JEPA intervene on what the predictor can see; we intervene on what physically happens. From one saved simulator state we run the dynamics under an action $a$ and under a reference action $a_{\varnothing}$, and train the model to predict the difference $Δz=z^{a}-z^{a_{\varnothing}}$ between the two latent futures. The resulting objective, Do-JEPA, has an effect loss, a support loss (where the action enters), a propagation loss (where its effect travels) and invariance losses (what must not change). In a synthetic system with object-aligned variables, support supervision finds the directly intervened object in 99.95% of test cases, where a sparse action mask sends the action to a nuisance slot in every case, and response-onset supervision recovers the ring-shaped propagation graph (edge AUROC 0.975 vs. 0.624). From pixels, the effect loss beats a control trained on exactly the same data: it lowers latent effect error by 28.4% on an end-to-end LeWM model and physical effect error by 13.5% when trained and tested on natural action sequences, and on three independently generated CausalWorld benchmarks it lowers responsive effect error by about 20% under physics shifts and the latent context sensitivity of predicted effects by 66%. Trained from scratch it costs factual accuracy; fine-tuning an existing model with it removes this cost. Together, these results show that intervening on the world, rather than on what the model sees, helps latent world models predict what their actions cause.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑