arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26239cs.RO

WALL-SS:通过下一尺度自回归扩展来缩放长视界世界模型

WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression

Maeve Zhang, Rain Sun, Xiang Wang, Cyril Zhang, Shalfun Li, Meng Cao, Howard Lu, Ethan Chen, Harry Jhou, KZ Zheng, Lights Shi, Regis Cheng, Lorenzin, Robert Wan… 展开作者

Maeve Zhang, Rain Sun, Xiang Wang, Cyril Zhang, Shalfun Li, Meng Cao, Howard Lu, Ethan Chen, Harry Jhou, KZ Zheng, Lights Shi, Regis Cheng, Lorenzin, Robert Wang, Victor Yao, Gody Li, Elise Mon, Yohann Tang, Ryan Yu, PS Zhang, Vincent Chen, Hang Su, Roy Gan, Hao Wang, Qian Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出WALL-SS世界模型,通过尺度自回归扩展实现动作可控的长视界机器人仿真,经实验验证其可提升动作跟随与轨迹精度,减少动作漂移和长视界不一致性。

中文摘要 AI 辅助

生成式世界模型为机器人提供了世界在交互下如何演化的预测模型,在仿真、规划、策略评估和机器人学习领域的潜力日益增长。除了片段级的未来预测,统一的生成式公式应能关联动作与结果、支持灵活视界和连续交互,并实现奖励驱动的优化。我们提出WALL-SS,一种通过尺度自回归扩展生成视觉未来的世界模型,可实现动作可控的长视界机器人仿真。WALL-SS将具身轨迹表示为时间上交错的观测与动作的因果序列,明确了依赖动作的状态转换,同时自然支持可变长度生成、通过可重用因果状态的流式扩展,以及通过序列概率的直接优化。为使该公式在长视界下有效,我们以由粗到细的方式生成每个未来观测,并在同一层级内开发三个互补组件:动作条件下一尺度预测注入与尺度对齐的动作表示,以改进动作-未来耦合并同时建模成功与失败行为;尺度压缩长视界记忆以精细分辨率保留近期交互,同时压缩远期观测与动作,而尺度级梦境强制增强对自生成上下文的鲁棒性;最后,在线策略对齐优化带动作跟随和长期一致性奖励的自回归视觉动态,同时保留预训练视觉分布。实验表明,WALL-SS提升了动作跟随与轨迹精度,在受限内存下支持连贯的分钟级流式回滚,且在线策略对齐始终有助于减少动作漂移和长视界不一致性。

英文摘要

Generative world models provide robots with predictive models of how the world evolves under interaction, with growing potential for simulation, planning, policy evaluation, and robot learning. Beyond clip-level future prediction, a unified generative formulation should relate actions to consequences, support flexible horizons and continuous interaction, and enable reward-driven optimization. We introduce WALL-SS, a world model that generates visual futures through Scale-wise autoregressive Scaling, enabling action-controllable and long-horizon robotic simulation. WALL-SS represents embodied trajectories as causal sequences of temporally interleaved observations and actions, making action-dependent state transitions explicit while naturally supporting variable-length generation, streaming extension through reusable causal states, and direct optimization through sequence probabilities. To make this formulation effective over long horizons, we generate each future observation in a coarse-to-fine manner and develop three complementary components within the same hierarchy. Action-conditioned next-scale prediction injects scale-aligned action representations to improve action-future coupling and model both successful and failed behaviors. Scale-compressed long-horizon memory retains recent interactions at fine resolution while compressing distant observations and actions, with scale-wise dream forcing enhancing robustness to self-generated context. Finally, on-policy alignment optimizes autoregressive visual dynamics with action-following and long-term consistency rewards while preserving the pretrained visual distribution. Experiments show that WALL-SS improves action following and trajectory accuracy, supports coherent minute-long streaming rollout under bounded memory, and consistently benefits from on-policy alignment in reducing action drift and long-horizon inconsistency.

发表机构

  • X Square Robot(X Square机器人)

机构由 AI 辅助整理,请以论文原文为准。

↑