SLIM-0.5B:学习机器人操作的动作基础预测隐变量
SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation
浏览论文内容
中文总结 AI 辅助
SLIM-0.5B是0.5B参数的紧凑机器人操作策略,通过自监督掩码轨迹预测学习动作基础预测隐变量,性能优于或匹配大型基线,参数量少、推理延迟低、内存占用小。
中文摘要 AI 辅助
视觉-语言-动作策略(Vision-language-action policies)依赖大型多模态主干网络,在每个控制步骤联合执行感知、语言条件化和动作生成。这些能力大多支持开放域语义,而连续机器人操作主要需要观测、动作及动作诱导的转移的紧凑表示。像素级世界模型是另一种途径,但预测与控制无关的视觉细节会造成不必要的开销。我们提出SLIM(自监督隐变量交互模型,Self-supervised Latent Interaction Model),一种参数规模为0.5B的紧凑隐变量交互策略。SLIM学习动作基础的预测隐变量,既捕获动作条件下的未来转移,也捕获解释观测变化的动作。SLIM通过自监督掩码轨迹预测学习这些表示,结合动作重构与未来隐变量预测。紧凑的混合Transformer(Mixture-of-Transformers,MoT)主干网络对观测隐变量与动作标记间的交互进行建模。所得策略经流匹配(flow matching)训练,用于语言条件化动作生成。在模拟基准测试与真实世界评估中,SLIM参数量更少、无需额外具身预训练、推理延迟更低、GPU内存使用显著更少,且性能与代表性大规模VLA及世界-动作模型基线相当或更优。
英文摘要
Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every control step. Much of this capacity supports open-domain semantics, whereas continuous robot manipulation primarily requires compact representations of observations, actions, and the transitions induced by actions. Pixel-level world models provide another route, but predicting visual details irrelevant to control can be unnecessarily expensive. We propose SLIM (Self-supervised Latent Interaction Model), a compact 0.5B-parameter latent interaction policy. SLIM learns action-grounded predictive latents that capture both action-conditioned future transitions and the actions that explain observed changes. SLIM learns these representations through self-supervised masked trajectory prediction, combining action reconstruction with future-latent prediction. A compact Mixture-of-Transformers (MoT) backbone models interactions between observation latents and action tokens. The resulting policy is trained with flow matching for language-conditioned action generation. Across simulation benchmarks and real-world evaluation, SLIM matches or exceeds representative large-scale VLA and world-action-model baselines with fewer parameters, no additional embodied pretraining, lower inference latency, and substantially lower GPU memory usage.
发表机构
- Fudan University(复旦大学)
- Beijing Academy of Artificial Intelligence(北京人工智能研究院)
- Tsinghua University(清华大学)
- Renmin University of China(中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。