PileBelief:面向交互驱动世界建模的持久物理状态
PileBelief: Persistent Physical State for Interaction-Driven World Modeling
- Tsinghua University(清华大学)
- Massachusetts Institute of Technology(麻省理工学院)
- Tsing-AI(Shanghai) Technology Co., Ltd(清智(上海)科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出PileBelief,一种交互驱动的持久世界模型,通过结合物理先验与记忆机制,在部分可观测挖掘中降低预测误差和动作选择遗憾,实现多步预测与候选动作排序。
AI中文摘要:
世界模型使机器人能够在执行动作之前预判动作后果。这一能力在挖掘作业中尤为宝贵,因为每一次铲取都会重塑地形并影响后续动作。然而,局部观测无法完全揭示底层的支撑条件和物料状况。我们提出了PileBelief,一种面向部分可观测挖掘场景的交互驱动持久世界模型,它能够保留超出可见表面的物理证据。该模型将观测条件化的物理先验与面向世界的变形记忆和物理响应记忆相结合。动作对齐的读取和门控残差校正可细化地形变化和结果预测。在部署权重固定的情况下,已完成交互会更新实测信念,而假设动作则推进一个独立的想象状态。与仅依赖当前观测的基线相比,PileBelief将五步联合预测误差降低了10.8%,离线动作选择遗憾值降低了65.5%。在Newton/MPM和真实挖掘数据集上的实验进一步证明了其在地形变化和铲斗体积预测方面的改进。我们的方法使得即使在底层土壤状态未知的情况下,也能从局部观测实现多步预测和候选动作排序。这些结果将持久物理信念确定为机器人持续重塑环境中世界模型的一种有效表示。
英文摘要:
World models allow robots to anticipate action consequences before execution. This capability is especially valuable in excavation, where each scoop reshapes the terrain and affects subsequent actions. Local observations, however, cannot fully reveal the underlying support and material conditions. We present PileBelief, an interaction-driven persistent world model for partially observed excavation that retains physical evidence beyond the visible surface. It combines an observation-conditioned physical prior with world-addressed deformation memory and physical-response memory. Action-aligned reads and gated residual corrections refine terrain-change and outcome predictions. With deployment weights fixed, completed interactions update measured belief, while hypothetical actions advance a separate imagined state. Compared with a current-observation-only baseline, PileBelief reduces five-step joint prediction error by 10.8% and offline action-selection regret by 65.5%. Experiments on Newton/MPM and real excavation datasets further demonstrate improved terrain-change and bucket-volume prediction. Our method enables multi-step prediction and candidate-action ranking from local observations, even when the underlying soil state is unknown. These results identify persistent physical belief as a useful representation for world models of environments that robots continually reshape.