AI 中文总结
提出GIFT框架,通过零初始化卷积将生成的目标图像注入预训练VLA模型,实现单周期微调即在SIMPLER和LIBERO上分别提升6.0%、13.4%和4.7%的性能。
AI 中文摘要
与仅依赖初始观察和语言指令相比,使用生成模型预测目标图像作为高层视觉引导,可以显著增强视觉-语言-动作(VLA)模型的鲁棒性。然而,由于高昂的计算训练成本,大多数现有基础模型尚未系统性地纳入目标图像条件。为此,我们提出了目标注入微调(GIFT),一个轻量且高效的微调框架,能够无缝地将生成的目标图像集成到多个具有代表性的预训练VLA模型中。我们的方法通过零初始化卷积将目标图像特征引入观察中,该卷积从零开始逐步增长参数,并在微调过程中防止有害噪声干扰预训练策略。随着训练的进行,目标信息逐渐被纳入,从而在不破坏模型稳定性的情况下实现高效的目标理解。我们进一步引入了一种精细的图像编辑方法,从初始观察和任务指令中生成语义和视觉一致的目标图像。实验表明,目标感知的VLA模型在各项任务中取得了显著的性能提升:仅经过一个周期的微调,GIFT在两个SIMPLER设置上分别比基础模型高出6.0%和13.4%,在LIBERO上高出4.7%,展示了其高效性和有效性。
英文摘要
Compared with relying solely on initial observations and language instructions, predicting goal images with generative models as high-level visual guidance can significantly enhance the robustness of Vision-Language-Action (VLA) models. However, most existing foundation models have not systematically incorporated goal image conditioning due to the high computational training cost. To this end, we propose Goal-Injected Fine-Tuning (GIFT), a lightweight and efficient fine-tuning framework that seamlessly integrates generated goal images into multiple representative pretrained VLA models. Our approach introduces goal image features into observations via a zero-initialized convolution which progressively grows parameters from zero and prevents harmful noise from disrupting the pretrained policy during fine-tuning. As training proceeds, goal information is gradually incorporated, enabling efficient goal understanding without disrupting model stability. We further introduce a refined image editing method to generate semantically and visually consistent goal images from initial observations and task instructions. Experiments show that goal-aware VLA models achieve substantial performance gains across tasks: with only a single epoch of fine-tuning, GIFT outperforms the base model by 6.0% and 13.4% on two SIMPLER settings, and by 4.7% on LIBERO, demonstrating both efficiency and effectiveness.
Comments20 pages, 12 figures, accepted by EMNLP2026