XS-VLA:将粗粒度空间蒸馏与潜在流匹配相结合用于轻量级机器人控制
XS-VLA: Teaching Tiny Vision-Language-Action Models with Spatial Supervision and Demonstration Conditioning
浏览论文内容
中文总结 AI 辅助
提出XS-VLA框架,先通过微调从Qwen3-VL-4B向SmolVLM2-0.25B骨干网蒸馏空间语义知识,再用其增强骨干网训练潜在流匹配策略,提升轻量级机器人控制性能,在LIBERO基准测试中表现出色。
中文摘要 AI 辅助
大型视觉语言模型计算成本高限制实时机器人控制,轻量级模型有‘空间盲目性’。训练视觉语言动作模型时人类示范多样会降低策略性能。为此提出XS-VLA,经粗粒度空间描述微调蒸馏知识,结合潜在流匹配策略,在LIBERO基准测试中表现优异,消融实验验证其有效性。
英文摘要
How can richer training supervision improve robot control while keeping the deployed policy compact? We present XS-VLA, a staged training framework using teacher-derived spatial labels and demonstration-conditioned action learning. Coarse-Grained Spatial Distillation (CSD) initializes the backbone through an auxiliary region-label task. Latent Flow Matching (LFM) then conditions an action-space velocity field on a demonstration latent, using KL regularization while jointly optimizing the backbone and action modules. The deployed policy contains 243.99M parameters and operates without the teacher or posterior encoder. XS-VLA achieves 90.25% average LIBERO success in each of two training seeds, compared with 86.00% for a SmolVLA-256M base trained under our settings. Ablations examine both training stages through matched image pretraining and Huber/MSE controls. On three Mobile ALOHA tasks, average strict success increases from 21.7% to 65.0%. These results demonstrate the control utility of auxiliary representation initialization and regularized demonstration-conditioned flow learning for compact VLA~policies.
发表机构
- Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
- National College for Excellent Engineers, Beihang University(北京航空航天大学卓越工程师学院)
- Wuxi Dexteroushands Robotic Technology Co.(无锡灵犀机器人技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。