发表机构
Drexel University; Virginia Tech; Amazon Store Foundation AI (SFAI)(德雷塞尔大学; 弗吉尼亚理工大学; 亚马逊商店基础人工智能(SFAI))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对VLA策略中可供性头注入问题,提出停止梯度保护主干并采用主动初始化的残差桥接,在LIBERO上达到96.2%,媲美复杂三专家模型,并揭示初始化是关键及桥接的承重作用。
AI 中文摘要
密集可供性监督是视觉-语言-行动(VLA)策略的一种有吸引力的辅助信号,然而,天真地共同训练一个可供性头会严重损害指令遵循能力。我们在LIBERO基准上对如何将这样的头接入现代VLA进行了受控研究。我们的方案通过停止梯度读取主干网络,并通过一个学习到的桥接层将中间头特征重新注入动作专家。停止梯度是一个先决条件:让可供性梯度到达主干网络会使策略性能降至低于无头基线(85.5%对比93.1%)。在保护主干网络的前提下,对注入拓扑(拼接对比残差)和桥接初始化(零初始化对比随机初始化)进行同预算的2*2消融实验表明,初始化是主导因素。最佳布线,即主动初始化的残差桥接,达到了96.2%的性能,与更为精细的三专家AffordanceVLA(95.8%)相当,且额外参数不到1%。两项探针实验解释了其机制:将真实可供性作为输入会损害性能,而推理时置零表明,惰性桥接仅作为训练时的正则化器,而主动桥接则成为承重组件。
英文摘要
Dense affordance supervision is an appealing auxiliary signal for vision-language-action (VLA) policies, yet naively co-training an affordance head can severely damage instruction following. We present a controlled study of how to wire such a head into a modern VLA on the LIBERO benchmark. Our recipe reads the backbone through a stop-gradient and re-injects an intermediate head feature into the action expert via a learned bridge. The stop-gradient is a precondition: letting affordance gradients reach the backbone drops the policy below the headless base (85.5% vs. 93.1%). With the backbone protected, a same-budget 2*2 ablation over injection topology (concatenation vs. residual) and bridge initialization (zero vs. random) shows initialization is the dominant lever. The best wiring, an actively initialized residual bridge, reaches 96.2%, matching the far more elaborate three-expert AffordanceVLA (95.8%) with under 1% extra parameters. Two probes explain the mechanism: ground-truth affordances fed as an input hurt, and inference-time zeroing shows a lazy bridge acts only as a training-time regularizer while an active bridge becomes load-bearing.
Comments8 pages, 4 figures, 2 tables