发表机构
New York University; Carnegie Mellon University; Columbia University(纽约大学; 卡内基梅隆大学; 哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过可控实验揭示,隐式世界模型的目标函数结构决定其获取的物理参数,输入与预测目标分别限制和决定可获知的物理量,额外数据仅优化已获取的参数。
AI 中文摘要
隐式世界模型的核心前提是,预测未来的任务会迫使表示学习器内化环境的物理规律。训练后的隐式表示究竟包含哪些物理量,又由什么因素决定?我们通过在POKEWORLD环境中开展可控干预实验来解答这一问题:该交互式环境中视觉外观完全相同的物体,实则隐藏着质量、阻力和接触刚度三类物理参数。我们采用一种带证书门控的协议,先对每个参数是否可从原始观测中恢复进行认证,再测量其是否进入隐式表示;若结果为阴性,则可归因于目标函数而非环境本身。由此得到的可识别性图谱包含两类组织机制和一个边界:输入会限制可获知的内容,而预测目标会决定需保留的内容。刚度仅在需要预测触摸时才会进入隐式表示(此时R²=0.50,而当同一信号仅作为输入融合时,R²=-0.02);在单步预测场景下,仅视觉输入的隐式表示会丢弃即使是完全可见的物体状态。阻力则构成了边界:它的可恢复性证书为0.89,但在我们测试的所有确定性预测目标下,其表现都稳定在0.13附近,而同一主干网络上的监督头可达到0.45。在感知坐标下读数缓慢且为比值型的参数,不在这些目标函数的获取范围内。在RH20T数据集上,跨缩放曲线的输入-目标因子设计在两个机器人和4258个 episodes 上复现了上述两类机制:缺少信息或预测压力的臂在五倍数据范围内保持平稳,仅完整多模态目标能迫使模型超越持续性基线,且保留的增益随数据规模增长。研究表明,目标函数结构决定了隐式表示会获取哪些物理参数,额外数据仅能改善模型已获取的参数。
英文摘要
A central premise of latent world models is that predicting the future encourages representations to internalize the physics of their environment. We ask which physical quantities are accessible in learned latent states, how this depends on training, and how those quantities relate to the model's predictions. We present PokeWorld, a simulated environment in which a robot finger pushes objects whose mass, drag, and contact stiffness vary across episodes while remaining visually identical. We first measure parameter recovery from raw observation sequences, then train matched action-conditioned world models. Prediction targets can strongly shape latent content. For example, contact stiffness becomes decodable when touch is predicted, while providing touch only as an input does not. Longer-horizon prediction improves estimation of position and velocity from learned representations relative to single-step prediction. Drag is recoverable from raw observations, but has weak linear readout from the learned latent states. Yet the models' glide forecasts depend systematically on drag. Experiments on RH20T reproduce the same input-target dependence on real multimodal robot data. Together, these results show that observations, prediction targets, and prediction horizons shape both which physical quantities are accessible in latent states and how they affect future predictions.