AI 中文总结
本研究提出轻量级世界-动作模型LiLa-WAM,通过联合未来状态预测与动作生成的紧凑隐式空间设计实现端到端单GPU训练,结合视觉转换令牌完成任务指定,在RoboTwin等任务上取得90.48%的成功率,提升了机器人操控的效率与实用性。
AI 中文摘要
世界-动作建模已成为机器人控制领域极具潜力的范式,它使模型不仅能对观测做出反应,还能预测场景的演变。然而,现有世界-动作模型(WAMs)往往会产生大量计算开销:像素空间方法会将大量容量分配给与控制无直接关联的视觉细节,而部分隐式空间方法则需要多阶段训练来构建推理空间,由此产生的训练成本使得这类方法难以在有限计算预算下完成训练。本研究提出LiLa-WAM,这是一种轻量级世界-动作模型,可在紧凑的隐式空间中对未来进行推理,且能在单块24GB GPU上进行端到端训练。其核心设计是由未来状态预测与动作生成共同塑造的紧凑隐式推理空间,在保持模型轻量化的同时仍与控制高度适配。针对任务指定,本研究进一步提出视觉转换令牌(VTT),这是一种无语言的任务表示方法,将每个任务编码为视觉特征空间中的一个方向。在RoboTwin 2.0、LIBERO及真实机器人任务上开展的实验证明了LiLa-WAM的有效性:在单GPU训练下,其在50项RoboTwin任务中达到了90.48%的成功率。
英文摘要
World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticipate how a scene will evolve. However, existing WAMs often incur substantial computational overhead. Pixel-space methods often allocate substantial capacity to visual details that may not be directly relevant to control, while some latent-space methods require multi-stage training to construct the reasoning space. The resulting training cost can make such methods difficult to train under modest computational budgets. In this work, we propose LiLa-WAM, a lightweight world-action model that reasons about the future in a compact latent space and can be trained end-to-end on a single 24GB GPU. Its core design is a compact latent reasoning space jointly shaped by future-state prediction and action generation, which keeps the model lightweight while remaining well aligned with control. For task specification, we further propose the Visual Transition Token(VTT), a language-free task representation that encodes each task as a direction in visual feature space. Experiments on RoboTwin~2.0, LIBERO, and real-robot tasks demonstrate LiLa-WAM's effectiveness, achieving 90.48\% success across 50 RoboTwin tasks with single-GPU training.