VLA-ZO:视觉-语言-动作模型的快速零阶自适应
VLA-ZO: Fast Zeroth-Order Adaptation for Vision-Language-Action Models
浏览论文内容
中文总结 AI 辅助
针对VLA模型部署时分布偏移,提出VLA-ZO框架,通过冻结视觉-语言前缀并重用条件状态,将ZO自适应时间减少25.59-32.54倍,并将任务成功率提升至63.58%。
中文摘要 AI 辅助
将视觉-语言-动作(VLA)模型适配到部署时的分布偏移对于可靠的机器人操作至关重要,但传统的一阶自适应可能会超出面向推理的部署平台的内存预算。零阶(ZO)优化提供了一种仅前向的替代方案,其内存开销与推理相当,但准确的梯度估计需要大量扰动查询,使得朴素ZO对于大型VLA模型而言慢得难以接受。我们提出了VLA-ZO,一个利用VLA计算结构实现快速ZO自适应的框架。通过将自适应限制在动作侧,VLA-ZO保持昂贵的视觉-语言前缀冻结,并在扰动查询和优化器步骤之间重用其条件状态,同时调度感知的预取隐藏了状态传输开销。在LIBERO相机视角偏移上,相对于基线ZO,VLA-ZO在q=16时将端到端自适应时间减少了25.59倍,在q=64时减少了32.54倍,同时将平均任务成功率从无自适应时的48.27%分别提高到58.17%和63.58%。这些结果表明,使ZO更快可以使更大的查询预算变得可行,为在部署平台上实现资源高效的VLA自适应提供了一条有前景的路径。
英文摘要
Adapting vision-language-action (VLA) models to deployment-time distribution shifts is important for reliable robotic operation, but conventional first-order adaptation can exceed the memory budget of inference-oriented deployment platforms. Zeroth-order (ZO) optimization offers a forward-only alternative with inference-level memory, but accurate gradient estimation requires many perturbation queries, making naive ZO prohibitively slow for large VLA models. We present VLA-ZO, a framework for fast ZO adaptation that exploits the structure of VLA computation. By confining adaptation to the action side, VLA-ZO keeps the expensive vision-language prefix frozen and reuses its conditioning states across perturbation queries and optimizer steps, while schedule-aware prefetching hides state-transfer overhead. On LIBERO camera-viewpoint shifts, VLA-ZO reduces end-to-end adaptation time by 25.59$\times$ at $q=16$ and 32.54$\times$ at $q=64$ relative to baseline ZO, while improving average task success from 48.27% without adaptation to 58.17% and 63.58%, respectively. These results show that making ZO faster can make larger query budgets practical, providing a promising path toward resource-efficient VLA adaptation on deployment platforms.
发表机构
- UNIST(蔚山科学技术院)
机构由 AI 辅助整理,请以论文原文为准。