arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RealtimeWAM:一步异步世界动作模型

RealtimeWAM: One-Step Asynchronous World Action Models

Chengtao Lv, Jinyang Du, Shuyi Feng, Yang Yong, Shiqiao Gu, Shunzi Yang, Ruihao Gong, Shen Ren, Tianwei Zhang, Wenya Wang

arXiv 2610.06617首次发表:更新:

发表机构

Nanyang Technological University; Beihang University; Sensetime; Continental Automotive Singapore(南洋理工大学; 北京航空航天大学; 商汤科技; 大陆汽车新加坡)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出RealtimeWAM,通过一步动作生成和异步推理解决WAM推理瓶颈,采用TACD和CEWP技术,在多个基准上保持近无损性能并实现约25倍加速。

AI 中文摘要

世界动作模型(WAMs)将视频生成骨干网络的视觉表示整合进来,以指导动作预测。近期的高效WAMs采用混合变换器(MoT)架构,并仅计算一次视频表示以供动作专家复用。然而,专家内部迭代(即多步动作去噪)和专家间等待(即视频和动作专家的顺序执行)仍然限制了推理效率。为此,我们提出了RealtimeWAM,一种具有一步动作生成和异步推理的极其高效的WAM变体,以解决这两个瓶颈。为减少专家内部迭代,我们提出了教师锚定一致性蒸馏(TACD)来解决局部-全局误差差距:仅低的局部一致性误差并不能保证最终动作的准确性。TACD通过冻结教师的多步展开终点的显式监督来补充局部一致性,从而实现准确的一步动作生成。此外,我们提出了跨专家波前流水线(CEWP)以消除不必要的专家级等待。它通过块级共享视频KV缓存来重叠两个专家,仅在相应的动作注意力消费该缓存之前立即同步。跨多个基准(如LIBERO、LIBERO-Plus和RoboTwin)和模型变体(如Fast-WAM和Faster-WAM)的大量实验证明了RealtimeWAM的优越性。值得注意的是,RealtimeWAM在这些基准上保持了近乎无损的性能(即性能下降<1%),同时提供了显著的端到端加速(如在H100上约25倍)。我们的代码和检查点可通过此链接获取。

英文摘要

World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action expert. However, intra-expert iteration (\ie, multi-step action denoising) and inter-expert waiting (\ie, sequential execution of the video and action experts) still limit inference efficiency. To this end, we present RealtimeWAM, an extremely efficient WAM variant with one-step action generation and asynchronous inference, addressing these two bottlenecks. To reduce intra-expert iteration, we propose Teacher-Anchored Consistency Distillation (TACD) to address a local-global error gap: low local consistency error alone does not guarantee accurate final actions. TACD supplements local consistency with explicit supervision from the frozen teacher's multi-step rollout endpoint, enabling accurate one-step action generation. Additionally, we propose Cross-Expert Wavefront Pipelining (CEWP) to eliminate unnecessary expert-level waiting. It overlaps the two experts through block-wise sharing of the video KV cache, synchronizing only immediately before the corresponding action attention consumes it. Extensive experiments across diverse benchmarks (\eg, LIBERO, LIBERO-Plus and RoboTwin) and model variants (\eg, Fast-WAM and Faster-WAM) demonstrate the superiority of RealtimeWAM. Notably, RealtimeWAM maintains near-lossless performance (\ie, $<1\%$ drop) across these benchmarks while delivering significant end-to-end speedup (\eg, $\sim25\times$ on H100). Our code and checkpoints are available via this \href{https://github.com/ModelTC/LightX2V/tree/main/examples/realtimewam}{link}.

CommentsThe code and checkpoints are available at $\href{https://github.com/ModelTC/LightX2V/tree/main/examples/realtimewam}{\text{this https URL}}$

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑