LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
LatBot: 从大规模物体操作视频中提炼通用潜在动作用于视觉-语言-动作模型
机构 * Institute of Microelectronics, Chinese Academy of Sciences(中国科学院微电子研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; Microsoft Research(微软研究院)
AI总结 LatBot通过整合动作预测和潜在动作分解,提升视觉-语言-动作模型在现实世界和模拟环境中的泛化与迁移能力。
Comments Project Page: https://mm-robot.github.io/distill_latent_action/