发表机构
Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Yinwang Intelligent Technology Co., Ltd.; Beijing Institute of Technology; Wuhan AI Research(中国科学院自动化研究所; 中国科学院大学人工智能学院; 银湾智能科技有限公司; 北京理工大学; 武汉人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对世界动作模型的预测-审议差距,提出HarnessWAM框架,通过双时间尺度反馈环等机制,在两个基准数据集上取得最优任务成功率,扩展了WAMs的具身任务执行能力。
AI 中文摘要
世界动作模型(World Action Models, WAMs)联合学习环境动态与机器人动作,为具身控制引入物理演化先验。然而,有限时域的预测与动作生成无法满足复杂具身任务需求,这类任务需要全局规划、跨阶段状态维护、执行验证及故障恢复,我们将这种不匹配称为WAMs的预测-审议差距。为解决该差距,我们提出HarnessWAM,一种面向WAMs的智能体框架。HarnessWAM采用基于视觉语言模型的任务管理器,以维持基于证据的场景信念与结构化任务图;能力条件可执行空间投影进一步将开放式语义计划约束为满足任务依赖、具身状态约束及底层WAM能力边界的原子技能序列。执行期间,HarnessWAM通过事件驱动的双时间尺度反馈环运作:轻量进度估计器持续提供高频执行证据,任务管理器在关键里程碑处结合当前观测、任务状态与交互历史进行审议,以决定推进任务、获取额外观测、修订计划或启动局部恢复。该机制使机器人能在子任务失败后恢复状态,且不丢弃已获取的场景知识即可恢复执行。HarnessWAM在RoboMemArena上实现了59.6%的全任务成功率与69.9%的子任务成功率,在RoboCerebra Ideal上实现了23.7%的成功率。这些结果表明,模型外部的结构化状态维护与闭环智能体决策可有效将WAMs的局部控制能力扩展为可规划、可验证且可恢复的具身任务执行能力。
英文摘要
World Action Models (WAMs) jointly learn environmental dynamics and robot actions, introducing priors over physical evolution into embodied control. However, finite-horizon prediction and action generation are insufficient for complex embodied tasks that require global planning, cross-stage state maintenance, execution verification, and failure recovery. We refer to this mismatch as the prediction-deliberation gap of WAMs. To address this gap, we propose HarnessWAM, an agentic framework for WAMs. HarnessWAM employs a vision-language-model-based Task Manager to maintain an evidence-grounded scene belief and a structured task graph. A capability-conditioned executable-space projection further constrains open-ended semantic plans into sequences of atomic skills that satisfy task dependencies, embodiment-state constraints, and the capability boundary of the underlying WAM. During execution, HarnessWAM operates through an event-driven, dual-timescale feedback loop: a lightweight progress estimator continuously provides high-frequency execution evidence, while the Task Manager deliberates at salient milestones by jointly considering the current observation, task state, and interaction history to determine whether to advance the task, acquire additional observations, revise the plan, or initiate local recovery. This mechanism enables the robot to recover its state after a subtask failure and resume execution without discarding previously acquired scene knowledge. HarnessWAM achieves state-of-the-art full-task and subtask success rates of 59.6% and 69.9% on RoboMemArena, and an SR of 23.7% on RoboCerebra Ideal. These results demonstrate that model-external structured state maintenance and closed-loop agentic decision making can effectively extend the local control capabilities of WAMs into embodied task execution that is plannable, verifiable, and recoverable.