发表机构
Technische Hochschule Nürnberg Georg Simon Ohm; Siemens AG; Fraunhofer Institute for Integrated Circuits (IIS)(纽伦堡乔治·西蒙·欧姆应用技术大学; 西门子股份公司; 弗劳恩霍夫集成电路研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究工业机器人操作中VLA模型的LoRA微调效率,通过在四个精密装配任务上评估,发现特定LoRA配置不比FFT差,r = 32时性能饱和,均匀分配即可,冻结VLM等会降性能,该方法可减少VRAM且无性能损失。
AI 中文摘要
在工业硬件上部署数十亿参数的视觉语言动作(VLA)模型需要微调以弥合实体差距。全量微调(FFT)具有最大可塑性,但需要数据中心级GPU。我们对基于流匹配的VLA模型$\pi_0$进行了低秩自适应(LoRA)的系统研究,在使用UR5e机器人操纵器的四个精密装配任务上进行评估。研究了不同LoRA秩、分配策略和组件冻结消融情况,发现FFT在某些LoRA配置下无显著优势。性能在r = 32时饱和,视觉语言模型(VLM)骨干和动作专家的均匀分配就足够。冻结VLM或限制视觉编码器为LoRA会显著降低性能。结果表明r = 32且视觉编码器全量微调的LoRA是实用方法,可减少静态峰值VRAM。
英文摘要
Deploying billion-parameter Vision-Language-Action (VLA) models on industrial hardware requires fine-tuning to bridge the embodiment gap. Full Fine-Tuning (FFT) provides maximal plasticity but requires data centre-grade GPUs. We present a systematic study of Low-Rank Adaptation (LoRA) for $π_0$, a flow-matching VLA, evaluated on four precision assembly tasks with a UR5e robotic manipulator. Across a sweep of LoRA ranks (r=8 to 256), allocation strategies, and component-freezing ablations, we find no statistically significant advantage of FFT over certain LoRA configurations. Performance saturates at r=32, and uniform allocation across the Vision-Language-Model (VLM) backbone and action expert proves sufficient. Freezing the VLM or restricting the vision encoder to LoRA significantly degrades performance, indicating that embodiment adaptation requires both semantic and visual plasticity. These results suggest that LoRA at r=32 with full vision encoder fine-tuning is a practical approach, reducing static peak VRAM from 36.2 to 10.8 GiB (parameters and optimizer states, activation memory excluded) without detectable performance loss.
Comments12 pages, 5 figures, 3 tables. Accepted at the International Conference on Artificial Neural Networks (ICANN 2026); to appear in Springer LNCS