arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过离线隐藏状态蒸馏恢复激进剪枝的视觉-语言-动作模型

Recovering Aggressively Pruned Vision-Language-Action Models with Offline Hidden-State Distillation

Chiyoung Kim, Sanghyuk Roy Choi, Minhyeok Lee

arXiv 2609.19579首次发表:更新:

发表机构

Chung-Ang University(中央大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出离线隐藏状态蒸馏方法,无需在线交互即可恢复激进剪枝的VLA模型,在OpenVLA-OFT和CogACT上显著提升成功率,并实现更快的机载推理与更低内存占用。

AI 中文摘要

视觉-语言-动作(VLA)模型使机器人能够遵循语言指令,但其数十亿参数的语言骨干是其在机器人硬件上运行的主要障碍。结构化剪枝可缩减该骨干,从OpenVLA-OFT中移除其63%的参数会使LIBERO-Long的成功率从93.2%骤降至0.8%。近期一种方法通过监督微调后接强化学习来恢复此类模型,这需要在线回放和数百个GPU小时。我们完全离线地恢复了大部分丢失的成功率。宽度剪枝收窄了块,但将残差流保持在其原始尺寸,因此教师和学生隐藏状态具有相同形状,可直接匹配,无需投影器。针对在一次教师前向传播中构建的缓存进行训练,可在约8个GPU小时内将缩减63%的学生模型提升至与教师相差3.5个百分点以内。对九个比例的扫描定位了恢复目标开始起作用的临界点。在OpenVLA-OFT上,缩减比例高达45%时,两者无显著差异。此后,隐藏状态蒸馏在63%至87%之间额外带来+2.1至+4.5个百分点,在CogACT上从63%起带来+9.4至+22.1个百分点。在CogACT的81%缩减比例下,三倍的恢复预算将蒸馏学生模型与教师的差距平均缩小至3.9个百分点,而监督恢复仍落后20个百分点以上。在相同压缩比下,宽度剪枝产生更高的成功率,而深度剪枝产生更低的延迟。在6自由度机械臂上,缩减72%的蒸馏学生模型达到77.5%的成功率,而监督恢复为59.5%,在机载运行速度比教师快2.23倍,且内存使用减少62%。

英文摘要

Vision-language-action (VLA) models let robots follow language instructions, but their language backbones of several billion parameters are the main obstacle to running them on robot hardware. Structured pruning reduces that backbone, and removing 63% of it from OpenVLA-OFT drops LIBERO-Long success from 93.2% to 0.8%. A recent approach restores such a model with supervised fine-tuning followed by reinforcement learning, which needs online rollouts and hundreds of GPU-hours. We recover most of the lost success entirely offline. Width pruning narrows the blocks but keeps the residual stream at its original size, so teacher and student hidden states have the same shape and are matched directly, without a projector. Training against a cache built in one teacher pass lifts the 63%-reduced student to within 3.5 points of the teacher in about 8 GPU-hours. A sweep over nine ratios locates where the recovery objective starts to matter. Up to 45% reduction the two do not differ significantly on OpenVLA-OFT. Hidden-state distillation then adds +2.1 to +4.5 points there between 63% and 87%, and +9.4 to +22.1 points on CogACT from 63% onward. At 81% on CogACT, a tripled recovery budget narrows the distilled student's gap to the teacher to 3.9 points on average, while supervised recovery stays more than 20 points below. At matched compression, width pruning yields higher success and depth pruning lower latency. On a 6-DoF manipulator, the distilled student at 72% reduction reaches 77.5% success against 59.5% for supervised recovery, runs 2.23x faster on-board than the teacher, and uses 62% less memory.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑