arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05483cs.ROcs.AI

FLEX-WAM:用于长时程想象与规划的灵活块因果世界-动作模型

FLEX-WAM: Flexible Block-Causal World-Action Models for Long-Horizon Imagination and Planning

R. Khorrambakht, Joseph Amigo, Félix Lebel, Leon Seetoo, Jean Ponce, Zhenzhen Li, Ludovic Righetti

首次发表
浏览论文内容

中文总结 AI 辅助

FLEX-WAM提出一种灵活高效的块因果世界-动作模型,支持长时程想象与规划,通过平衡梯度贡献和FD弹性提升动作响应性,在模拟和真实任务中实现稳定推演与实时部署。

中文摘要 AI 辅助

世界-动作模型(WAMs)有望成为一个统一模型,能够预测动作条件下的未来状态、生成可行动作,并支持在想象中进行规划。然而,现有的联合视频-动作模型通常采用计算量大、固定时域的主干网络,不适合流式推理和稳定的长时程开环推演。我们提出了FLEX-WAM,一种灵活高效的块因果世界-动作模型,用于统一的仿真和策略推理。FLEX-WAM支持可变长度上下文和非因果预测时域,以及逐帧或逐块的无限自回归生成。其块因果、可缓存KV的架构结合了轴向注意力和块状扩散强制,实现了高效的实时推演和部署时的延迟-吞吐量权衡,无需重新训练。然而,联合训练可能产生对命令动作响应较弱的不切实际的未来。我们通过平衡状态和动作流匹配在状态-动作扩散噪声网格上的梯度贡献,并使用前向动力学(FD)弹性(一种高效的训练时动作响应性代理)来调节世界模型采样,从而解决这一失败模式。在模拟和真实世界数据集上,FLEX-WAM实现了优越的多步预测质量和延迟,同时为数千步生成稳定的联合状态-动作推演。作为MCTS中的联合动作提议器和模拟器,它完全在想象中解决了长时程PushT任务和所有五个OGBench Puzzle-4x4任务。在基于OpenArm的双臂机器人上,单个检查点同时作为游戏策略和预期结果预测器,实现模型-现实不匹配的实时识别和收集,以促进未来的自我改进。

英文摘要

World--action models (WAMs) promise a unified model that predicts action-conditioned futures, generates feasible actions, and supports planning in imagination. However, existing joint video--action models often use computationally heavy, fixed-horizon backbones ill-suited to streaming inference and stable long-horizon open-loop rollouts. We introduce FLEX-WAM, a Flexible and Efficient Block-Causal World--Action Model for unified simulation and policy inference. FLEX-WAM supports variable-length contexts and non-causal prediction horizons, as well as infinite autoregressive generation frame by frame or block by block. Its block-causal, KV-cacheable architecture combines axial attention and blockwise diffusion forcing to enable efficient real-time rollout and deployment-time latency--throughput tradeoffs without retraining. Joint training can nevertheless produce plausible futures that weakly respond to commanded actions. We address this failure mode by balancing state and action flow-matching gradient contributions across the state--action diffusion-noise grid and regulating world-model sampling using Forward-Dynamics (FD) elasticity, an efficient training-time proxy for action responsiveness. Across simulated and real-world datasets, FLEX-WAM achieves superior multi-step prediction quality and latency while producing stable joint state--action rollouts for thousands of steps. As a joint action proposer and simulator within MCTS, it solves long-horizon PushT and all five OGBench Puzzle-4x4 tasks entirely in imagination. On a bimanual OpenArm-based robot, a single checkpoint jointly serves as a play policy and expected-outcome predictor, enabling real-time identification and collection of model--reality mismatches for future self-improvement.

发表机构

  • Center for Robotics and Embodied Intelligence (CREO), New York University(纽约大学机器人与具身智能中心)
  • Courant Institute of Mathematical Sciences and Center for Data Science, New York University(纽约大学库朗数学科学研究所和数据科学中心)
  • Ecole normale supérieure - PSL(巴黎高等师范学校-巴黎文理研究大学)
  • NVIDIA(英伟达)
  • Artificial and Natural Intelligence Toulouse Institute (ANITI)(图卢兹人工与自然智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑