具身化GPT-5.1:世界模型的证据?
Embodied GPT-5.1: Evidence of a World Model?
浏览论文内容
中文总结 AI 辅助
研究探索无实体化训练的GPT-5.1能否控制物理移动机器人,通过低分辨率图像和离散动作集执行任务。它展现出空间推理等新兴能力,但也有效率和感知局限。结果挑战传统观点,促使深入研究大语言模型中物理理解相关问题。
中文摘要 AI 辅助
本探索性研究考察了大型多模态语言模型GPT-5.1,在没有先前实体化、未在模拟环境中训练且未接触过感觉运动经验的情况下,能否作为物理移动机器人的高级控制器。仅使用低分辨率第一人称图像和离散动作集,让该模型执行导航和目标导向行为,如定位和接触目标玩具。在多次试验中,GPT-5.1展现出空间推理和物理理解等新兴能力,包括物体离开相机画面后保持短期记忆、推断自身动作的物理后果等,但也存在效率低下和感知局限。结果表明,即便未接受任何与实体化相关训练,GPT-5.1在具身化环境中展现出类似世界模型的行为迹象,这一发现挑战了认知科学和机器人学中认为物理身体是发展此类智能必要前提的长期观点,促使对大语言模型中物理理解的出现、局限和鲁棒性进行更深入研究。
英文摘要
This exploratory study examines whether a large multimodal language model, GPT-5.1, can serve as the high-level controller of a physical mobile robot despite having no prior embodiment, no training in simulated environments, and no exposure to sensorimotor experience. Using only low-resolution first-person images and a discrete action set, the model was tasked with navigation and object-directed behaviors such as locating and contacting a target toy. Across multiple trials, GPT-5.1 demonstrated emergent capabilities that suggest elements of spatial reasoning and physical understanding. These included maintaining short-term memory of object locations after they left the camera frame, inferring the physical consequences of its own movements, and executing coherent action sequences such as colliding with an object and reversing to visually verify the outcome. At the same time, the model displayed inefficiencies and perceptual limitations, including imprecise alignment strategies and occasional misidentification of distant distractors. Overall, the results indicate that GPT-5.1 exhibits signs of world-model-like behavior in an embodied setting, despite the absence of any embodiment-related training, a finding that challenges long-standing views in cognitive science and robotics which hold that a physical body is a necessary prerequisite for developing such forms of intelligence. The findings motivate deeper investigation into the emergence, limits, and robustness of physical understanding in large language models.