WALL-SS:通过下一尺度自回归扩展来缩放长视界世界模型
WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression
浏览论文内容
中文总结 AI 辅助
本研究提出WALL-SS世界模型,通过尺度自回归扩展实现动作可控的长视界机器人仿真,经实验验证其可提升动作跟随与轨迹精度,减少动作漂移和长视界不一致性。
中文摘要 AI 辅助
生成式世界模型为机器人提供了世界在交互下如何演化的预测模型,在仿真、规划、策略评估和机器人学习领域的潜力日益增长。除了片段级的未来预测,统一的生成式公式应能关联动作与结果、支持灵活视界和连续交互,并实现奖励驱动的优化。我们提出WALL-SS,一种通过尺度自回归扩展生成视觉未来的世界模型,可实现动作可控的长视界机器人仿真。WALL-SS将具身轨迹表示为时间上交错的观测与动作的因果序列,明确了依赖动作的状态转换,同时自然支持可变长度生成、通过可重用因果状态的流式扩展,以及通过序列概率的直接优化。为使该公式在长视界下有效,我们以由粗到细的方式生成每个未来观测,并在同一层级内开发三个互补组件:动作条件下一尺度预测注入与尺度对齐的动作表示,以改进动作-未来耦合并同时建模成功与失败行为;尺度压缩长视界记忆以精细分辨率保留近期交互,同时压缩远期观测与动作,而尺度级梦境强制增强对自生成上下文的鲁棒性;最后,在线策略对齐优化带动作跟随和长期一致性奖励的自回归视觉动态,同时保留预训练视觉分布。实验表明,WALL-SS提升了动作跟随与轨迹精度,在受限内存下支持连贯的分钟级流式回滚,且在线策略对齐始终有助于减少动作漂移和长视界不一致性。
英文摘要
Generative world models provide robots with predictive models of how the world evolves under interaction, with growing potential for simulation, planning, policy evaluation, and robot learning. Beyond clip-level future prediction, a unified generative formulation should relate actions to consequences, support flexible horizons and continuous interaction, and enable reward-driven optimization. We introduce WALL-SS, a world model that generates visual futures through Scale-wise autoregressive Scaling, enabling action-controllable and long-horizon robotic simulation. WALL-SS represents embodied trajectories as causal sequences of temporally interleaved observations and actions, making action-dependent state transitions explicit while naturally supporting variable-length generation, streaming extension through reusable causal states, and direct optimization through sequence probabilities. To make this formulation effective over long horizons, we generate each future observation in a coarse-to-fine manner and develop three complementary components within the same hierarchy. Action-conditioned next-scale prediction injects scale-aligned action representations to improve action-future coupling and model both successful and failed behaviors. Scale-compressed long-horizon memory retains recent interactions at fine resolution while compressing distant observations and actions, with scale-wise dream forcing enhancing robustness to self-generated context. Finally, on-policy alignment optimizes autoregressive visual dynamics with action-following and long-term consistency rewards while preserving the pretrained visual distribution. Experiments show that WALL-SS improves action following and trajectory accuracy, supports coherent minute-long streaming rollout under bounded memory, and consistently benefits from on-policy alignment in reducing action drift and long-horizon inconsistency.
发表机构
- X Square Robot(X Square机器人)
机构由 AI 辅助整理,请以论文原文为准。