arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对变化进行建模:面向以对象为中心的操作的稀疏残差世界模型

Modeling What Changes: Sparse, Residual World Models for Object-Centric Manipulation

Param Thakkar, Parsika Paresh Shah, Manisha Sushant Gote

arXiv 2609.02046首次发表:更新:

发表机构

Veermata Jijabai Technological Institute; Arizona State University; ZuiGO Private Limited(维尔玛塔·吉贾巴伊理工学院; 亚利桑那州立大学; ZuiGO私人有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出稀疏残差世界模型,通过显式建模对象变化,在以对象为中心的操作任务上实现更高预测准确率、更少参数量与更好迁移性,且能为规划器提供有效特征支持。

AI 中文摘要

整体式世界模型每一步都会预测整个下一状态,将算力浪费在重复预测场景中占多数的静态部分,进而引入误差。本文探究显式建模变化(即每个对象的变化门,加上仅对该门标记的对象进行扰动的残差增量头)是否是更有效、更可解释的物理预测与控制偏置。在从3个对象扩展到8个对象的MuJoCo桌面推压基准任务上,该稀疏/残差模型的下一状态姿态预测准确率是密集多层感知机的2.5至4.6倍,参数量却少8.6至11.1倍;其变化检测F1值维持在0.80至0.87,而密集基线模型已退化;该模型可跨对象数量迁移,无需重新训练(F1值保留率达99.4%),仅用四分之一的数据就能达到全数据约90%的准确率。在自回归滚动时,其误差累积远小于密集模型,在无运动基准附近保持稳定,而密集模型则发生漂移。最后,在基于采样的规划器中,仅用于预测的模型会失效(尽管真实模拟器神谕用同一规划器解决了该任务,证实规划器本身是可靠的),但该稀疏模型经特征化并针对规划器访问的状态训练后,开始具备规划能力(3个随机种子上的成功率为0.23±0.06),而密集整体模型在所有种子上均为0。对变化进行建模而非重复预测整个世界,是面向以对象为中心的物理AI的简单、有效偏置;代码、数据生成器及所有检查点将在发表后发布。

英文摘要

Monolithic world models predict the entire next state at every step, spending capacity re-predicting the static majority of a scene and injecting error into it. We ask whether explicitly modeling change (a per-object change gate plus a residual delta head that perturbs only the objects the gate flags) is a more effective and interpretable bias for physical prediction and control. On a MuJoCo tabletop pushing benchmark scaling from 3 to 8 objects, the sparse/residual model predicts next-state poses 2.5 to 4.6 times more accurately than a dense multilayer perceptron at 8.6 to 11.1 times fewer parameters, sustains change-detection F1 of 0.80 to 0.87 where the dense baseline is degenerate, transfers across object counts with zero retraining (99.4 percent F1 retention), and reaches about 90 percent of its full-data accuracy with a quarter of the data. In autoregressive rollout it compounds far less error, hugging the no-motion floor while the dense model drifts. Finally, inside a sampling-based planner, prediction-only models fail (though a true-simulator oracle solves the task with the identical planner, confirming the planner is sound), but once featurized and trained for the states a planner visits, the sparse model begins to plan (0.23 plus or minus 0.06 success over three seeds) while the dense monolith stays at zero at every seed. Modeling what changes, rather than re-predicting the whole world, is a simple, effective bias for object-centric physical AI; code, data generators, and all checkpoints will be released upon publication.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑