arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09771cs.RO

SLIM-0.5B:学习机器人操作的动作基础预测隐变量

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

Jingkai Wang, Zihan Tang, Gu Zhang, Mingyu Cao, Jiapeng Chen, Jingjiao Zhao, Xiansheng Chen, Pengwei Wang, Lemao Liu, Dejing Dou

首次发表
浏览论文内容

中文总结 AI 辅助

SLIM-0.5B是0.5B参数的紧凑机器人操作策略,通过自监督掩码轨迹预测学习动作基础预测隐变量,性能优于或匹配大型基线,参数量少、推理延迟低、内存占用小。

中文摘要 AI 辅助

视觉-语言-动作策略(Vision-language-action policies)依赖大型多模态主干网络,在每个控制步骤联合执行感知、语言条件化和动作生成。这些能力大多支持开放域语义,而连续机器人操作主要需要观测、动作及动作诱导的转移的紧凑表示。像素级世界模型是另一种途径,但预测与控制无关的视觉细节会造成不必要的开销。我们提出SLIM(自监督隐变量交互模型,Self-supervised Latent Interaction Model),一种参数规模为0.5B的紧凑隐变量交互策略。SLIM学习动作基础的预测隐变量,既捕获动作条件下的未来转移,也捕获解释观测变化的动作。SLIM通过自监督掩码轨迹预测学习这些表示,结合动作重构与未来隐变量预测。紧凑的混合Transformer(Mixture-of-Transformers,MoT)主干网络对观测隐变量与动作标记间的交互进行建模。所得策略经流匹配(flow matching)训练,用于语言条件化动作生成。在模拟基准测试与真实世界评估中,SLIM参数量更少、无需额外具身预训练、推理延迟更低、GPU内存使用显著更少,且性能与代表性大规模VLA及世界-动作模型基线相当或更优。

英文摘要

Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every control step. Much of this capacity supports open-domain semantics, whereas continuous robot manipulation primarily requires compact representations of observations, actions, and the transitions induced by actions. Pixel-level world models provide another route, but predicting visual details irrelevant to control can be unnecessarily expensive. We propose SLIM (Self-supervised Latent Interaction Model), a compact 0.5B-parameter latent interaction policy. SLIM learns action-grounded predictive latents that capture both action-conditioned future transitions and the actions that explain observed changes. SLIM learns these representations through self-supervised masked trajectory prediction, combining action reconstruction with future-latent prediction. A compact Mixture-of-Transformers (MoT) backbone models interactions between observation latents and action tokens. The resulting policy is trained with flow matching for language-conditioned action generation. Across simulation benchmarks and real-world evaluation, SLIM matches or exceeds representative large-scale VLA and world-action-model baselines with fewer parameters, no additional embodied pretraining, lower inference latency, and substantially lower GPU memory usage.

发表机构

  • Fudan University(复旦大学)
  • Beijing Academy of Artificial Intelligence(北京人工智能研究院)
  • Tsinghua University(清华大学)
  • Renmin University of China(中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑