arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18084cs.ROcs.CVcs.LG

并非所有层都需要微调:诊断与引导视觉-语言-动作模型中的适配

Not All Layers Need Tuning: Diagnosing and Directing Adaptation in Vision-Language-Action Models

发表机构卡内基梅隆大学
查看机构详情
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Shahram Najam Syed, Arthur Jakobsson, Prayuj Sachdev, Jeffrey Ichnowski

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种诊断与分配流水线,通过无微调估计VLA模型各区域适配成本,并据此分配变秩LoRA适配器,在多个架构与任务上以极低参数开销达到或超越全微调性能。

中文摘要 AI 辅助

针对新的部署环境微调视觉-语言-动作(VLA)模型成本高昂,然而大多数方法对所有网络区域采用统一容量的适配器,仿佛每个区域都需要同等程度的调整。本文在五个架构多样的VLA模型(OpenVLA-OFT、$\pi_0$、SmolVLA、DTP、Octo;参数量93M-7B)上检验了这一假设。通过区域隔离微调下的归一化参数位移来度量每区域的适配成本,揭示了一个适配谱系:在所有五个架构中,外观变化使成本集中于视觉编码器,指令变化集中于语言主干,而新物体变化则同时集中于视觉编码器和动作头。为利用这一结构,我们提出了一条观察、诊断、分配与适配的流水线。仅凭十个无标注目标观测且无需微调,诊断模块通过结合无参考梯度与蒙特卡洛Dropout信号,以及与缓存源参考的居中核对齐得分,来估计每区域成本;分配模块在参数预算下将估计转换为变秩LoRA适配器,并冻结校准良好的区域;随后标准LoRA微调训练所得适配器。该诊断在每次部署中对区域排序的中位Spearman相关系数为0.91,且在LIBERO和CALVIN上测试的每个预算下,其分配结果均达到或超过均匀LoRA。在物理xArm-7机械臂上,该流水线在指令措辞变化下以全微调0.04%的可训练参数达到全微调性能;在五个未经重训练评估的留出场景中,它领先所有基线,30次 rollout 中成功11-23次,而最强参数高效基线在相同或更大预算下为8-18次,全微调为2-11次。这些结果表明,VLA中的适配成本具有足够的结构性,可在微调开始前进行度量。

英文摘要

Fine-tuning a Vision-Language-Action (VLA) model for a new deployment environment is expensive, yet most methods apply uniform-capacity adapters to every network region as if every region requires equal adjustment. This paper tests that assumption on five architecturally diverse VLAs (OpenVLA-OFT, $π_0$, SmolVLA, DTP, Octo; 93M-7B parameters). Measuring per-region adaptation cost as normalized parameter displacement under region-isolated fine-tuning reveals an adaptation spectrum in which appearance shifts concentrate cost in the vision encoder, instruction shifts in the language backbone, and novel-object shifts in the vision encoder together with the action head, across all five architectures. To exploit this structure, we introduce a pipeline that observes, diagnoses, allocates, and adapts. From ten unlabeled target observations and without fine-tuning, the diagnostic estimates per-region cost by combining reference-free gradient and Monte Carlo Dropout signals with a Centered Kernel Alignment score against a cached source reference; the allocator converts the estimates into variable-rank LoRA adapters under a parameter budget and freezes well-calibrated regions; and standard LoRA fine-tuning trains the resulting adapters. The diagnostic ranks regions within each deployment at a median Spearman of 0.91, and the allocation matches or exceeds uniform LoRA at every budget we tested on LIBERO and CALVIN. On a physical xArm-7, the pipeline matches full fine-tuning under an instruction-wording shift with 0.04% of its trainable parameters, and on five held-out scenes evaluated without retraining it leads every baseline, with 11-23 successes of 30 rollouts against 8-18 for the strongest parameter-efficient baseline at equal or larger budgets and 2-11 for full fine-tuning. These results suggest that adaptation cost in VLAs is structured enough to measure before fine-tuning begins.

补充信息

↑