发表机构
Tsinghua University; YuanxingGuangnian Robotics; Nanjing University(清华大学; 元星广年机器人; 南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出RealtimeWAM,一种无需训练的通用框架,通过并行执行与自适应计算(选择性缓存与残差复用)降低世界动作模型推理延迟,在多种架构上实现8.9-10.7倍加速,同时保持或提升成功率。
AI 中文摘要
世界动作模型(WAMs)将视觉动力学建模与动作生成相结合,但其高推理延迟限制了响应式机器人控制。近期的工作通过去除测试时的显式未来视频生成来加速推理,如FastWAM,该方法需要专门定制的架构设计。更通用的缓存策略利用特征冗余,但仅靠冗余无法捕捉闭环控制中不断变化的计算需求。为解决这些挑战,我们提出了RealtimeWAM,一个通用的、无需训练的框架,该框架协调并行执行与自适应计算,以在多种WAM架构上实现低延迟推理。我们利用层间依赖关系将观测处理与预测重叠。然而,并发分支仍会竞争GPU资源,限制了并行执行的收益。因此,我们通过选择性复用调整整个流水线的计算:在视觉稳定区域缓存观测特征,并复用Transformer残差,同时为小的预测调整保留额外的细化。我们在FastWAM和OpenWAM上,针对RoboTwin、LIBERO和LIBERO-Plus评估了RealtimeWAM。在RTX 4090上,实测平均推理延迟分别为24.09毫秒和63.09毫秒,对应平均加速比分别为8.90倍和10.67倍。平均成功率分别为82.75%和87.41%,与原生推理相差0.02和0.53个百分点。在五个真实世界任务中,RealtimeWAM在FastWAM和OpenWAM上相比原生推理将平均成功率分别提高了17.2和37.2个百分点。
英文摘要
World Action Models (WAMs) combine visual dynamics modeling with action generation, but their high inference latency limits responsive robot control. Recent efforts accelerate inference by removing explicit future-video generation at test time, as in FastWAM, an approach that requires a specially tailored architectural design. More general caching strategies exploit feature redundancy, but redundancy alone does not capture the changing computational demands of closed-loop control. To address these challenges, we present RealtimeWAM, a general, training-free framework that coordinates parallel execution with adaptive computation for low-latency inference across diverse WAM architectures. We exploit layerwise dependencies to overlap observation processing with prediction. However, concurrent branches still compete for GPU resources, limiting the benefit of parallel execution. We therefore adapt computation throughout the pipeline through selective reuse, caching observation features in visually stable regions and reusing Transformer residuals while reserving additional refinement for small predicted adjustments. We evaluate RealtimeWAM on FastWAM and OpenWAM across RoboTwin, LIBERO, and LIBERO-Plus. On an RTX 4090, measured mean inference latencies are 24.09 and 63.09 ms, corresponding to average speedups of 8.90$\times$ and 10.67$\times$. Average success rates are 82.75% and 87.41%, respectively, within 0.02 and 0.53 percentage points of native inference. Across five real-world tasks, RealtimeWAM improves average success rates over native inference by 17.2 and 37.2 percentage points on FastWAM and OpenWAM, respectively.
Comments24 pages, including appendix. Project page: https://anonymous.4open.science/w/realtimewam/