arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Faster-WAM:面向鲁棒世界动作模型的高效推理时未来条件化方法

Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models

Weiheng Zhao, Haoyi Jiang, Xin Shi, Liu Liu, Fan Huang, Zhizhong Su, Wei Sui, Xinggang Wang

arXiv 2608.04404首次发表:更新:

AI 中文总结

本文提出Faster-WAM,通过稀疏未来条件化框架等技术解决现有WAMs的性能效率矛盾,在多个机器人操控基准上实现更优权衡,提升了分布外泛化能力与鲁棒性。

AI 中文摘要

世界动作模型(World Action Models, WAMs)通过学习当前观测之外的环境演化规律来提升机器人操控能力。然而现有方法存在一个根本困境:Joint-WAMs在推理阶段保留未来感知表征,但会产生高昂的计算成本;而高效替代方案在推理阶段移除未来建模,可能会损失时序推理带来的鲁棒性优势。本研究重新审视WAMs中未来表征的作用,发现推理时的未来条件化对于分布偏移下的泛化能力至关重要。基于该观察,本文提出Faster-WAM,一种高效的未来条件化WAM,它在保留未来表征的同时避免了昂贵的视频-动作交互。Faster-WAM引入稀疏未来条件化框架,该框架仅计算一次未来表征,并在整个动作去噪过程中有选择性地复用这些表征。具体而言,本文提出SparseMoT,用网络紧凑子集阶段的选择性视频-动作交互替代普遍存在的逐层融合;还提出Interval KV-Fusion,用于聚合多深度未来表征且不增加注意力复杂度。实验表明,Faster-WAM在性能与效率的权衡上显著优于现有WAMs:在分布外LIBERO-Plus基准上,相比Fast-WAM,其成功率从49.14%提升至73.57%,同时运行速度比Joint-WAM快2.21倍;此外,它在LIBERO和RoboTwin 2.0上达到了最先进的性能,并在真实世界操控中展现出强鲁棒性。

英文摘要

World Action Models (WAMs) improve robot manipulation by learning how the environment evolves beyond the current observation. However, existing approaches face a fundamental dilemma: Joint-WAMs preserve future-aware representations during inference but incur prohibitive computation costs, while efficient alternatives remove future modeling at inference time and may lose the robustness benefits of temporal reasoning. In this work, we revisit the role of future representations in WAMs and show that inference-time future conditioning is critical for generalization under distribution shifts. This observation motivates Faster-WAM, an efficient future-conditioning WAM that preserves future representations while avoiding expensive video-action interaction. Faster-WAM introduces a sparse future-conditioning framework that computes future representations once and selectively reuses them throughout action denoising. Specifically, we propose SparseMoT to replace ubiquitous layer-wise fusion with selective video-action interaction at a compact subset of network stages, and Interval KV-Fusion to aggregate multi-depth future representations without increasing attention complexity. Experiments demonstrate that Faster-WAM achieves a substantially better performance-efficiency trade-off than existing WAMs. On the out-of-distribution LIBERO-Plus benchmark, Faster-WAM improves success rate from 49.14% to 73.57% compared with Fast-WAM, while running 2.21$\times$ faster than Joint-WAM. It further achieves state-of-the-art performance on LIBERO and RoboTwin 2.0, while demonstrating strong robustness in real-world manipulation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑