arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Sparse-WAM:通过动作引导的稀疏想象加速世界动作模型

Sparse-WAM: Accelerating World Action Models via Action-Guided Sparse Imagination

Xinling Xie, Haodong Wang, Jiazhi Mi, Zhiming Liu, Zicong Hong, Xiaoyi Pang, Qianli Liu, Yangjia Hu, Ying Chen, Zhengyang Yan, Song Guo

arXiv 2609.38984首次发表:更新:

发表机构

NJU; HKUST; HIT; EPFL(南京大学; 香港科技大学; 哈尔滨工业大学; 瑞士洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Sparse-WAM,一种无需训练的动作引导稀疏想象框架,通过选择性地处理未来帧令牌并引入Pilot引擎,在LIBERO和RoboLab-120上分别实现约2.0倍和1.8倍的推理加速,同时保持任务性能。

AI 中文摘要

世界动作模型(WAMs)利用预训练的视频模型,通过联合预测未来的视觉状态和动作来提高机器人控制的泛化能力。这一能力带来了巨大的推理成本,因为在去噪过程中需要反复处理密集的未来帧令牌。先前的方法通过优先考虑视觉保真度的令牌剪枝来解决这一问题,以降低视频扩散模型中的去噪成本。然而,这些方法在WAMs的联合去噪过程中,并未利用动作相关性来决定保留哪些未来帧令牌。在本文中,我们提出了Sparse-WAM,一种无需训练的框架,用于动作引导的稀疏想象,通过选择性地处理未来帧令牌来加速WAM推理。我们观察到,在连续的去噪步骤之间,从动作令牌到未来帧令牌(动作到未来注意力)的注意力空间分布存在显著重叠,尽管未来表示持续更新。基于此,我们开发了动作引导的令牌选择,以保留帧特定的动作相关区域以及跨帧上下文。然而,简单的实现可能会产生注意力评分和令牌打包的开销,抵消剪枝带来的计算节省。因此,我们引入了Pilot,一种高效的引擎,通过轻量级评分和跨步骤重用令牌选择来减少稀疏推理的开销。在带有FastWAM-Joint的LIBERO和带有Cosmos 3 Edge的RoboLab-120上,Sparse-WAM在NVIDIA RTX 4090上相比密集急切推理分别实现了约2.0倍和1.8倍的推理加速,同时基本保持了任务性能。

英文摘要

World-action models (WAMs) leverage pretrained video models to improve generalization in robot control by jointly predicting future visual states and actions. This capability comes at a substantial inference cost, as dense future-frame tokens are repeatedly processed during denoising. Prior methods address this by token pruning that prioritizes visual fidelity to reduce denoising costs in video diffusion models. However, these methods do not use action relevance to determine which future-frame tokens to retain during joint denoising in WAMs. In this paper, we propose Sparse-WAM, a training-free framework for action-guided sparse imagination that selectively processes future-frame tokens to accelerate WAM inference. We observe substantial overlap in the spatial distribution of attention from action tokens to future-frame tokens (action-to-future attention) between consecutive denoising steps, despite continued updates to the future representations. Motivated by this, we develop Action-Guided Token Selection to retain frame-specific action-relevant regions together with cross-frame context. However, a naive implementation can incur attention-scoring and token-packing overhead that offsets the computational savings from pruning. We therefore introduce Pilot, an efficient engine that reduces sparse inference overhead through lightweight scoring and cross-step reuse of token selections. On LIBERO with FastWAM-Joint and RoboLab-120 with Cosmos 3 Edge, Sparse-WAM achieves inference speedups of approximately $2.0\times$ and $1.8\times$, respectively, over dense eager inference on an NVIDIA RTX 4090, while largely preserving task performance.

CommentsXinling Xie and Haodong Wang contributed equally to this work

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑