VLA 的层之谜:VLA 潜在接口的信息论分析
The Layer Mystery of VLA: An Information-Theoretical Analysis of VLA Latent Interface
- University of California, Berkeley(加州大学伯克利分校)
- California Institute of Technology(加州理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过信息论分析揭示VLA策略中潜在接口的层选择问题,发现多层融合优势有限,而InfoNCE作为层质量代理可高效选择最佳层,显著降低计算成本和选择遗憾。
AI中文摘要:
视觉-语言-动作(VLA)策略通过潜在接口将预训练的视觉-语言骨干网络与动作头连接起来,但该接口应暴露哪些骨干网络层仍不清楚。我们针对三个预训练模型和两个操作基准(LIBERO 和 CALVIN)研究了冻结骨干网络的单层选择和多层融合,每种配置使用三个策略训练种子。在三种融合机制和三种层子集策略下,54 种配置中有 47 种的表现不如观察到的最佳单层策略。我们的统计分析进一步证实,融合的优势非常有限。然而,最佳层在不同骨干网络和基准之间差异很大,这使得层选择至关重要,而穷举策略扫描代价高昂。我们进一步推导了所提出的用于动作条件 InfoNCE 和动作预测的信息瓶颈目标之间的重加权等价性,从而将 InfoNCE 作为层质量的代理指标。在评估的四种代理指标中,经验上 InfoNCE 与策略成功之间的一致性正相关最强。选择具有最高 InfoNCE 分数的层所需的 GPU 计算量比穷举策略扫描少 9 到 33 倍,并将六个设置中的平均选择遗憾从最深层的 17.89 个百分点降低到 3.71 个百分点。其平均遗憾接近使用所有六个 oracle 扫描事后优化的固定层启发式方法所达到的 3.28-3.50 个百分点,而无需在选择过程中进行闭环评估。
英文摘要:
Vision-language-action (VLA) policies connect a pretrained vision-language backbone to an action head through a latent interface, but which backbone layers this interface should expose remains unclear. We study single-layer selection and multi-layer fusion for frozen backbones across three pretrained models and two manipulation benchmarks, LIBERO and CALVIN, with three policy-training seeds per configuration. Across three fusion mechanisms and three layer-subset strategies, 47 of 54 configurations underperform the best observed single-layer policy. Our stastical analysis further confirms that fusion's advantage is very limited. However, the best layer varies substantially across backbones and benchmarks, making layer selection consequential and exhaustive policy sweeps expensive. We further derive a reweighting equivalence between the proposed information-bottleneck objectives for action-conditioned InfoNCE and action prediction, motivating InfoNCE as a proxy for layer quality. Empirically, InfoNCE provides the most consistent positive association with policy success among four evaluated proxies. Selecting the layer with the highest InfoNCE score requires 9-33 times less GPU compute than exhaustive policy sweeps and reduces mean selection regret from 17.89 percentage points for deepest-layer selection to 3.71 points across six settings. Its mean regret is close to the 3.28-3.50 points achieved by fixed-layer heuristics optimized retrospectively using all six oracle sweeps, without requiring closed-loop evaluations during selection.