WAM-Cache:用于高效世界动作模型的 stale 受限 KV 重用
WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models
浏览论文内容
中文总结 AI 辅助
WAM-Cache 是一种无训练框架,通过结合动作专家交叉注意力与视觉潜在惊喜选择刷新集,在抑制累积误差的同时,将视频 DiT 预填充 FLOPs 降低 32%-42%,且精度损失极小。
中文摘要 AI 辅助
世界动作模型(World Action Models, WAMs)通过将动作专家模型基于预训练视频扩散 Transformer(Diffusion Transformer, DiT)的表示进行条件化,实现通用机器人操作。在闭环控制中,视频 DiT 会在每个 chunk 运行,将当前观测编码为分层键值(Key-Value, KV)对,供动作专家查询。这种预填充操作占据了每个 chunk 的计算成本,而现有的无训练加速方法仍将其保持为全密度计算。我们提出 WAM-Cache,这是一种无训练框架,可在多个 chunk 间保留分层键值表示,仅重新计算稀疏的刷新 token 集。关键在于,我们发现仅刷新视觉漂移 token 的直观启发式方法,即使使用能预测真实 KV 漂移的 oracle,其性能也远低于全密度基线。下游动作精度由动作专家的注意力位置决定,而非移动的内容。因此,WAM-Cache 通过结合动作专家的交叉注意力与视觉潜在惊喜来选择刷新集,并辅以严格的年龄边界以抑制累积误差。在 Fast-WAM 上,WAM-Cache 在 RoboTwin 2.0、LIBERO 及真实世界实验中,将视频 DiT 预填充的 FLOPs 降低了 32%-42%,同时在仿真中与全密度策略的精度差距保持在 0.7-1.8 个百分点,在真实机器人上则为 2.5 个百分点。
英文摘要
World Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT). In closed-loop control, the video DiT runs at every chunk to encode the current observation into layerwise key-value (KV) pairs that the action expert queries. This prefill dominates the per-chunk computational cost, yet existing training-free accelerations leave it fully dense. We present WAM-Cache, a training-free framework that retains layerwise key-value representations across chunks and recomputes only a sparse refresh set of tokens. Crucially, we find that the intuitive heuristic of refreshing visually drifted tokens plateaus far below the dense baseline, even with an oracle predicting ground-truth KV drift. Downstream action accuracy is instead governed by where the action expert attends, not by what moved. WAM-Cache therefore selects the refresh set by uniting the action expert's cross-attention with visual latent surprise, complemented by a strict age bound that suppresses compounding error. On Fast-WAM, WAM-Cache cuts video DiT prefill FLOPs by 32-42% across RoboTwin 2.0, LIBERO, and real-world experiments, while staying within 0.7-1.8 percentage points of the dense policy in simulation and 2.5 points on a real robot.
发表机构
- Zhejiang University(浙江大学)
- Agency for Science, Technology and Research(科技研究局)
机构由 AI 辅助整理,请以论文原文为准。