arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27549cs.CV

代码即世界:用于物理推理的可执行世界表征的智能体式发现

Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning

  • MirroS

机构由 AI 辅助整理,请以论文原文为准。

Hanyang Wang, Yimo Cai, Weiliang Chen, Jiawei Chi, Haowen Sun, Qiyu Dai, Yi-Hsin Hung, Xingzhuo Guo, Jinshan Ren, Runmao Yao, Ziwei Liu, Mingsheng Long, Yueqi D… 展开作者

Hanyang Wang, Yimo Cai, Weiliang Chen, Jiawei Chi, Haowen Sun, Qiyu Dai, Yi-Hsin Hung, Xingzhuo Guo, Jinshan Ren, Runmao Yao, Ziwei Liu, Mingsheng Long, Yueqi Duan, Jun Gao, Jiangran Lyu, Fangfu Liu, Jialong Wu

AI总结:

本研究提出Code-as-World范式,通过智能体式发现循环构建可执行世界表征,将其用于为视觉-语言模型训练提供物理监督,在QuantiPhy数据集上取得最优性能且超越领先专有模型

AI中文摘要:

物理理解与推理依赖于形成紧凑且可泛化的世界表征。尽管现代视觉-语言模型能够识别并解释多样的物理事件,但它们往往缺乏对底层机制的显式表征,比如物体状态、物理参数以及支配动力学的规律,而这些机制是可靠地推理世界如何演化、对干预做出响应所必需的。在本研究中,我们提出了Code-as-World这一范式,该范式通过可执行世界表征来表征物理世界。通过将物理构成、动态演化以及视觉外观表达为可执行代码,Code-as-World为物理世界提供了一种紧凑、定量依据且可控的抽象。为了从多模态观测(如自然语言描述或现实世界视频)中构建此类表征,我们开发了一种受溯因推理启发的智能体式发现循环,其中智能体提出、执行、渲染、验证并迭代优化可执行世界假设。作为一项具体应用,我们利用经过验证的可执行世界为视觉-语言模型在定量物理推理任务上的训练提供可扩展的物理监督。实验表明,Code-as-World-VL在QuantiPhy数据集上达到了当前最优性能,并且超越了领先的专有模型,凸显了可执行世界表征作为物理智能可扩展基础的潜力。

英文摘要:

Physical understanding and reasoning depend on forming compact and generalizable representations of the world. While modern vision-language models can recognize and explain diverse physical events, they often lack explicit representations of the underlying mechanisms-such as object states, physical parameters, and governing dynamics-needed for reliably reasoning how the world evolves and responds to interventions. In this work, we introduce Code-as-World, a paradigm that represents physical worlds through executable world representations. By expressing physical composition, dynamic evolution, and visual appearance as executable code, Code-as-World provides a compact, quantitatively grounded, and controllable abstraction of the physical world. To construct such representations from multimodal observations, such as natural-language descriptions or real-world videos, we develop an agentic discovery loop inspired by abductive reasoning, where an agent proposes, executes, renders, verifies, and iteratively refines executable world hypotheses. As a concrete application, we use verified executable worlds to provide scalable physical supervision for training vision-language models on quantitative physical reasoning. Experiments show that Code-as-World-VL achieves state-of-the-art performance on QuantiPhy and surpasses leading proprietary models, highlighting the potential of executable world representations as a scalable foundation for physical intelligence.

补充信息

↑