发表机构
KAIST; Sungkyunkwan University(韩国科学技术院; 成均馆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FastOPD通过在线策略蒸馏将大规模VLA模型压缩为轻量级学生,结合流映射与自一致性目标,在LIBERO和RoboTwin上实现高成功率与低延迟,并成功部署于真实机器人。
AI 中文摘要
视觉-语言-动作(VLA)基础模型已快速扩展以增强操作性能和泛化能力,但这种扩展带来了高昂的计算成本,使得现实世界部署日益困难。现有方法通常通过设计更小的架构或减少基于流的策略中的迭代去噪步骤来缓解这一问题。在本工作中,我们提出FastOPD,一种从基础模型到轻量级VLA的框架,通过高效的在线策略蒸馏实现大规模VLA的实际部署。具体而言,FastOPD调整一个流映射用于单状态教师监督,并将其与自一致性目标相结合,构建一个学习教师动态的紧凑学生模型。此外,我们从理论上证明,最小化该目标可使蒸馏后的学生恢复与理想少步教师模型所诱导分布相当的数据分布。我们在模拟和真实世界实验中评估了FastOPD在多种基础策略上的表现。在LIBERO上,FastOPD仅用两步推理保留了π0.5性能的84%,推理延迟降低78.1%,同时在平均成功率上优于现有少步蒸馏基线。以LingBot-VLA为教师,FastOPD在RoboTwin 2.0上将单步成功率较基础学生提高了15.9个百分点。我们进一步展示了其应用于世界动作模型(WAM)的可行性,并在真实机器人上部署了从MolmoAct2蒸馏出的紧凑学生模型。
英文摘要
Vision-Language-Action (VLA) foundation models have scaled rapidly to enhance manipulation performance and generalizability, but this scaling incurs high computational costs that render real-world deployment increasingly challenging. Existing approaches typically mitigate this issue by designing smaller architectures or reducing the iterative denoising steps in flow-based policies. In this work, we propose FastOPD, a foundation-to-lightweight VLA framework that enables the practical deployment of large-scale VLAs through efficient on-policy distillation. Specifically, FastOPD adapts a flow map for single-state teacher supervision and combines it with a self-consistency objective to construct a compact student that learns the teacher dynamics. Furthermore, we theoretically demonstrate that minimizing this objective allows the distilled student to recover a distribution on par with that induced by an ideal few-step teacher model. We evaluate FastOPD across diverse foundation policies in simulation and real-world experiments. On LIBERO, FastOPD retains 84% of the performance of $π_{0.5}$ with only two inference steps, reducing inference latency by 78.1% while outperforming existing few-step distillation baselines in average success rate. With LingBot-VLA as the teacher, FastOPD improves the single-step success rate over the base student by 15.9 percentage points on RoboTwin 2.0. We further demonstrate its applicability to a World Action Model (WAM) and deploy a compact student distilled from MolmoAct2 on a real robot.
CommentsProject page: https://fastopd.github.io/