arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34608cs.ROcs.AI

高效世界动作模型推理与自适应中间状态

Efficient World Action Model Inference with Adaptive Intermediate States

发表机构厦门大学 · 清华大学人工智能产业研究院 · 北京大学计算机学院
另 2 家 · 查看机构详情
  • Xiamen University(厦门大学)
  • Institute for AI Industry Research, Tsinghua University(清华大学人工智能产业研究院)
  • School of Computer Science, Peking University(北京大学计算机学院)
  • Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

Zhinnan Liu, Haozhi Han, Ruge Zhang, Teng Ma, Tao Ma, Zheng Liu, Yifeng Chen, Yunquan Zhang, Ting Cao, Yunxin Liu, Kun Li

首次发表
浏览论文内容

中文总结 AI 辅助

提出WAMachine框架,通过轨迹重映射、观测重绑定和残差缩放调整推理状态,加速世界动作模型推理,在保持高成功率的同时显著降低延迟。

中文摘要 AI 辅助

世界动作模型(WAMs)通过联合建模动作与环境动态,实现面向未来的控制。然而,迭代式扩散或流推理会产生大量的去噪延迟。先前的推理状态为加速提供了自然的机会,但不断变化的规划上下文、观测和中间表示可能迅速使保留的状态失效。因此,保留有用的计算需要调整推理状态,而非直接复用。为此,我们提出WAMachine,一个无需训练的框架,通过保留并调整推理状态,在控制循环演化过程中实现高效且准确的延续。在闭环重规划中,轨迹重映射将上一次重规划的状态重新映射以初始化下一次重规划,减少冗余的轨迹生成。在去噪步骤中,观测重绑定在动作执行期间进行预期推理,并在一致性检查通过时将保留的去噪状态重新绑定到真实观测以继续推理,从而减少暴露于控制循环的延迟。在Transformer层中,残差缩放选择性地缩放保留的层状态,并在探针检查失败时通过中间层的完整计算刷新这些状态,从而减少重复的Transformer计算。在LIBERO和RoboTwin 2.0上对三种代表性WAM架构的评估表明,WAMachine在观测到动作的延迟上实现了1.47-3.05倍的加速,在每次重规划的GPU推理时间上实现了2.23-3.27倍的加速,同时保持了原生WAM任务成功率的96.69-99.54%。

英文摘要

World Action Models (WAMs) enable future-aware control by jointly modeling actions and environment dynamics. However, iterative diffusion or flow inference incurs substantial denoising latency. Prior inference state offers a natural opportunity for acceleration, yet changing planning contexts, observations, and intermediate representations can quickly render retained state stale. Preserving useful computation therefore requires adapting inference state rather than reusing it as-is. To this end, we present $\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}}$, a training-free framework that accelerates WAM inference by preserving and adapting inference state for efficient and accurate continuation as the control loop evolves. Across closed-loop replans, Trajectory Remapping remaps replan state from the preceding replan to initialize the next replan, reducing redundant trajectory generation. Across denoising steps, Observation Rebinding performs anticipatory inference during action execution and rebinds retained denoising state to the real observation for continuation when consistency checks pass, reducing latency exposed to the control loop. Across Transformer layers, Residual Rescaling selectively rescales retained layer state and refreshes it through full computation of the middle layers when probe checks fail, reducing repeated Transformer computation. Evaluations of three representative WAM architectures on LIBERO and RoboTwin 2.0 show that $\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}}$ achieves 1.47-3.05$\times$ speedups in observation-to-action latency and 2.23-3.27$\times$ speedups in GPU inference time per replan, while preserving 96.69-99.54% of native WAM task success.

↑