arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

潜在世界模型中的干预差距

The Intervention Gap in Latent World Models

Donna Vakalis

arXiv 2608.29998首次发表:更新:

发表机构

Mila – Quebec AI Institute; University of Montreal(米拉-魁北克人工智能研究所; 蒙特利尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现潜在世界模型存在干预差距,即模型开环转移与环境干预对任务变量的影响存在偏差,需以捕获优先的方式在模型原生接口直接审计干预保真度。

AI 中文摘要

规划时间的干预保真度是学习得到的世界模型的一种独特、可测量的属性,即模型自身的开环转移是否能像匹配的环境干预那样改变任务变量。在我们测试的场景中,该属性既未通过奖励拟合体现,也未通过以任务为锚定的训练得到保证。在发布的TD-MPC2检查点规模中,随着任务可观测变量上的算子误差诊断值增大,回合回报下降,而奖励预测误差保持较小且几乎平稳;在共享任务上,未使用任务信号训练的自监督世界模型,相比以任务为锚定的模型,能更好地保留该算子。随后,捕获门控匹配干预审计定位了失效点:在猎豹(Cheetah)任务中,三个LeWorldModel检查点能捕获当前任务查询并支持可解码的真实干预效果,但它们想象的五步效果比预测无效果更差,也比环境端点 oracle 更差,失效表现为任务方向旋转且增益过高,而非特征崩溃。这种严重模式具有条件性:五个PreJEPA种子在无此模式时仍保留相对于oracle的缺陷;手指旋转(Finger Spin)实验将该缺陷扩展到运动之外,不同种子的严重程度存在异质性;共享库效应几何既依赖候选也依赖支持。我们还测试了实践层面的问题:在DreamerV3中,后验分布而非其样本承载当前查询;集成分歧仅在训练支持附近对误差进行排名;冻结的支持感知评分会降低两种测试迁移方向上的保留误差排名,而原生分歧在两种方向上仍具有信息价值。我们得出结论,必须在模型的原生接口上以捕获优先的方式直接审计干预保真度。

英文摘要

Planning-time intervention fidelity is a distinct, measurable property of a learned world model: whether the model's own open-loop transitions move task variables the way matched environment interventions do. In the settings we test, it is neither revealed by reward fit nor ensured by task-anchored training. Across released TD-MPC2 checkpoint sizes, episode return falls as an operator-error diagnostic on task observables grows, while reward-prediction error stays small and nearly flat, and a self-supervised world model trained without task signal preserves the same operator substantially better than a task-anchored model on the shared task. A capture-gated matched-intervention audit then localizes what fails. On Cheetah, three LeWorldModel checkpoints capture the current task query and support decodable real intervention effects; however, their imagined five-step effects are worse than predicting no effect and worse than an environment-endpoint oracle. The failure is task-direction rotation with excess gain, not feature collapse. This severe pattern is conditional: five PreJEPA seeds retain an oracle-relative deficit without it, Finger Spin experiments extend the deficit beyond locomotion with heterogeneous severity across seeds, and shared-bank effect geometry is both candidate- and support-dependent. We also test practice-side questions. In DreamerV3 the posterior distribution, not its sample, carries the current query; ensemble disagreement ranks error only near training support; and a frozen support-aware score degrades held-out error ranking in both tested transfer directions while native disagreement remains informative in both. We conclude that intervention fidelity must be audited directly, capture-first, on the model's native interface.

Comments21 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑