arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37721cs.ROcs.CV

CogWAM:通过事件驱动接口将语义认知与世界动作建模对齐

CogWAM: Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces

发表机构地平线机器人 · 华中科技大学 · 中国科学技术大学
查看机构详情
  • Horizon Robotics(地平线机器人)
  • Huazhong University of Science and Technology(华中科技大学)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Sen Wang, Liu Liu, Xinjiang Wang, Zequn Chen, Haoyi Jiang, Taojun Ding, Tingyang Xiao, Zhizhong Su, Jie Wang, Sanping Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

CogWAM通过持久语义状态和事件驱动接口,将语义认知与世界动作建模对齐,实现任务级上下文持久化,在RoboDojo和BiCoord上取得领先性能,并减少真实世界双臂操作中的状态更新次数。

中文摘要 AI 辅助

机器人策略日益融入语义推理和未来世界预测,然而,将这些能力结合起来并不能保证局部预测和动作与任务进展保持一致。我们提出CogWAM,一种认知引导的世界动作模型,通过一个持久的语义状态(Semantic State)在任务推理与世界动作学习之间建立显式的语义接口,该状态存储已完成的任务事件和当前活动的子任务。CogWAM仅在观测表明发生语义转换时更新此状态,从而使任务级上下文能够在多个动作块之间持续存在。为了将语义上下文与物理预测和控制相连接,CogWAM采用进展条件化的WORLD和ACTION查询,选择性地提取与任务相关的信息用于未来世界预测和动作生成。在训练期间,语义状态为两个分支提供共享的任务进展上下文,而推理时则移除未来预测分支,直接从观测和维护的状态生成动作。我们进一步引入语义训练策略,以改进转换学习和闭环条件化。在没有额外机器人动作预训练的情况下,CogWAM在RoboDojo上达到15.56 / 11.70 %的Score/SR,并在BiCoord上取得最先进性能,同时真实世界实验展示了闭环双臂操作,其语义状态再生次数比逐步更新少16.4次。

英文摘要

Robot policies increasingly incorporate semantic reasoning and future-world prediction, yet combining these capabilities does not guarantee that local predictions and actions remain aligned with task progress. We introduce CogWAM, a cognition-guided world-action model that establishes an explicit semantic interface between task reasoning and world-action learning through a persistent Semantic State, which stores completed task events and the active subtask. CogWAM updates this state only when observations indicate semantic transitions, allowing task-level context to persist across multiple action chunks. To bridge semantic context with physical prediction and control, CogWAM employs progress-conditioned WORLD and ACTION queries that selectively extract task-relevant information for future-world prediction and action generation. During training, the Semantic State provides shared task-progress context for both branches, while inference removes the future-prediction branch and directly generates actions from observations and the maintained state. We further introduce semantic training strategies to improve transition learning and closed-loop conditioning. Without additional robot-action pretraining, CogWAM achieves 15.56 / 11.70 % Score/SR on RoboDojo and state-of-the-art performance on BiCoord, while real-world experiments demonstrate closed-loop dual-arm manipulation with 16.4 fewer Semantic State regenerations than step-wise updating.

↑