发表机构
Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对物理交互规划中实例槽无法明确实体任务角色的问题,提出SR-WM模型,通过功能角色绑定实现语义接口,在LIBERO模拟套件等评估中验证了其视觉动力学与规划决策的连接能力。
AI 中文摘要
用于物理交互的世界模型通常被训练来预测未来观测或潜在特征;然而,面向规划的模型必须回答一个根本不同的问题:候选动作是否会产生与任务一致的未来,同时保留必要的状态表示。标准的实例级对象槽仅识别存在的“什么”,而未指定每个实体在任务上下文中扮演的“什么角色”。为弥合这一差距,我们提出了语义丰富世界模型(Semantically Rich World Model, SR-WM),这是一种围绕五个功能角色构建的任务条件世界模型:夹爪、目标、任务目标、关系和环境。在SR-WM中,视觉实体编码器从预训练的补丁特征中提取软实体假设,允许分割掩码作为可选的提案先验,而不将其强制要求为必需的状态表示或推理输入。随后,角色绑定器将这些假设映射到特定任务的角色,而动作条件动力学模型则预测角色转换以及细粒度语义,包括抓取/接触、谓词建立、关系保留、固定装置状态和阶段信息。这种统一的角色状态为下游多候选动作生成、阶段感知重排序和违规感知后缀提供了基础。我们的综合评估协议涵盖所有四个LIBERO模拟套件、跨套件迁移、感知诊断和动作敏感性测试。总体而言,该公式将以对象为中心的预测转化为一个语义接口,将视觉动力学与面向规划的决策联系起来。
英文摘要
World models for physical interaction are typically trained to predict future observations or latent features; however, a planning-oriented model must answer a fundamentally different question: whether a candidate action produces a task consistent future while preserving essential relations. Monolithic state representations obscure the underlying entities, while standard instance-level object slots merely identify what is present without specifying what role each entity plays in the task context. To bridge this gap, we present the Semantically Rich World Model (SR-WM), a task-conditioned world model structured around five functional roles: gripper, target, goal, relation, and phase. Within SR-WM, a visual entity encoder extracts soft entity hypotheses from pretrained patch features, allowing segmentation masks to serve as optional proposal priors without mandating them as required state representations or inference inputs. A role binder subsequently maps these hypotheses to task-specific roles, while an action conditioned dynamics model predicts role transitions alongside fine-grained semantics, including grasp/contact, predicate establishment, relation preservation, fixture state, and phase change. Crucially, this unified role state grounds downstream multi-candidate action generation, stage-aware reranking, and violation-aware suffix resampling. Our comprehensive evaluation protocol spans all four LIBERO simulation suites, cross-suite transfer, perception diagnostics, and action sensitivity analysis. Ultimately, this formulation transforms object-centric prediction into a semantic interface linking visual dynamics with planning-oriented decision making