发表机构
College of Computer Science and Technology, National University of Defense Technology; School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences; Intelligent Game and Decision Lab (IGDL); School of Artificial Intelligence, Shanghai Jiao Tong University(国防科技大学计算机学院; 中国科学院大学电子电气与通信工程学院; 智能博弈与决策实验室; 上海交通大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出VIDEAS框架,从演示视频中蒸馏显式动作语义,结合先验引导仿真验证,并训练VIDEAS-WM世界模型,实现高层次具身动作推理的最先进性能。
AI 中文摘要
世界模型学习环境动态的内部表征以预测未来状态,使智能体无需物理交互即可优化行动计划。然而,开发真正内化潜在因果物理规律、以显式推理动作前置条件及后续状态转换的世界模型,仍是一个开放挑战。本文提出VIDEAS,一种数据蒸馏框架,将操作视频中的连续物理动态转化为基础模型的显式动作语义。具体而言,它把视觉演示分解为离散动作轨迹,并利用先进视觉-语言模型(VLM)提取封装动作前置条件与效果的结构化知识。为确保物理一致性,我们引入一种基于文本环境的先验引导轨迹仿真机制,严格验证所提取的知识。值得注意的是,我们纳入负轨迹以丰富知识完整性,并增强数据多样性以缓解认知偏差。此外,我们提出VIDEAS-WM,一个基于AgiBot-World数据集、由34K高质量样本训练而成的8B/9B参数语言世界模型套件。大量实验表明,VIDEAS-WM在高层次具身动作语义推理中达到最先进性能,展现出深刻的物理理解及对未见场景的稳健泛化能力。
英文摘要
World models learn internal representations of environment dynamics to predict future states, enabling agents to optimize action plans without physical interactions. However, developing world models that genuinely internalize underlying causal physical laws to explicitly reason about action preconditions and subsequent state transitions remains an open challenge. In this paper, we propose VIDEAS, a data distillation framework that transforms continuous physical dynamics from operational videos into explicit action semantics for foundation models. Specifically, it deconstructs visual demonstrations into discrete action trajectories and utilizes advanced vision-language models (VLMs) to extract structured knowledge encapsulating action preconditions and effects. To ensure physical consistency, we introduce a prior-guided trajectory simulation mechanism grounded within a text-based environment to rigorously validate the extracted knowledge. Notably, we incorporate negative trajectories to enrich knowledge completeness and enhance data diversity to mitigate cognitive bias. Furthermore, we present VIDEAS-WM, an 8B/9B-parameter suite of language-based world models trained on 34K high-quality samples derived from AgiBot-World dataset. Extensive experiments demonstrate that VIDEAS-WM establishes state-of-the-art performance in high-level embodied action semantic reasoning, exhibiting profound physical understanding and robust generalization across unseen scenarios.