arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.04171cs.ROcs.LG

XS-VLA:将粗粒度空间蒸馏与潜在流匹配相结合用于轻量级机器人控制

XS-VLA: Teaching Tiny Vision-Language-Action Models with Spatial Supervision and Demonstration Conditioning

Iok Tong Lei, Ying Jie Yap, Wei Huang, Qingchen Xie, Qianzhi Li, Yujie Zhang, Xiaolong Liu, Zhidong Deng

首次发表
浏览论文内容

中文总结 AI 辅助

提出XS-VLA框架,先通过微调从Qwen3-VL-4B向SmolVLM2-0.25B骨干网蒸馏空间语义知识,再用其增强骨干网训练潜在流匹配策略,提升轻量级机器人控制性能,在LIBERO基准测试中表现出色。

中文摘要 AI 辅助

大型视觉语言模型计算成本高限制实时机器人控制,轻量级模型有‘空间盲目性’。训练视觉语言动作模型时人类示范多样会降低策略性能。为此提出XS-VLA,经粗粒度空间描述微调蒸馏知识,结合潜在流匹配策略,在LIBERO基准测试中表现优异,消融实验验证其有效性。

英文摘要

How can richer training supervision improve robot control while keeping the deployed policy compact? We present XS-VLA, a staged training framework using teacher-derived spatial labels and demonstration-conditioned action learning. Coarse-Grained Spatial Distillation (CSD) initializes the backbone through an auxiliary region-label task. Latent Flow Matching (LFM) then conditions an action-space velocity field on a demonstration latent, using KL regularization while jointly optimizing the backbone and action modules. The deployed policy contains 243.99M parameters and operates without the teacher or posterior encoder. XS-VLA achieves 90.25% average LIBERO success in each of two training seeds, compared with 86.00% for a SmolVLA-256M base trained under our settings. Ablations examine both training stages through matched image pretraining and Huber/MSE controls. On three Mobile ALOHA tasks, average strict success increases from 21.7% to 65.0%. These results demonstrate the control utility of auxiliary representation initialization and regularized demonstration-conditioned flow learning for compact VLA~policies.

发表机构

  • Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
  • National College for Excellent Engineers, Beihang University(北京航空航天大学卓越工程师学院)
  • Wuxi Dexteroushands Robotic Technology Co.(无锡灵犀机器人技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑