发表机构
Shanghai Qizhi Institute; The Chinese University of Hong Kong; Hong Kong Embodied AI Lab; Tsinghua University; University of California, Berkeley(上海期智研究院; 香港中文大学; 香港具身人工智能实验室; 清华大学; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉-语言-动作模型中动作监督问题,提出动作QFormer,通过基于指令的查询重组多模态信息。在零样本模拟到真实导航中提升了任务成功率、动作生成正确性等,还改变动作监督塑造多模态表示的方式,为提高VLA性能提供新思路。
AI 中文摘要
视觉-语言-动作(VLA)模型中的动作监督通常被视为学习动作预测的下游目标。本文将其视为塑造继承多模态表示的力量。这种塑造有双重作用:对形成动作兼容表示是必要的,但直接应用于继承多模态路径时会破坏支持语言处理和对象基础的表示。为解决此问题,引入动作QFormer,它使用基于指令的查询在下游动作生成前将继承多模态信息重组为面向动作的表示。在零样本模拟到真实导航中,它提高了平均闭环任务成功率、固定指令动作生成正确性并减少分布外指令生成。进一步分析表明它改变了动作监督塑造继承多模态表示的方式。这些结果表明提高VLA性能不仅需要更强的预训练主干,还需要更好的方式来选择和组织继承多模态信息并控制其在动作监督下的塑造。
英文摘要
Action supervision in vision-language-action (VLA) models is often treated as a downstream objective for learning action prediction. In this paper, we study it instead as a force that shapes inherited multimodal representations. We show that this shaping has a dual effect: it is necessary for forming action-compatible representations, but when action supervision is applied too directly to the inherited multimodal pathway, it can also destabilize representations that support language-side processing and object grounding. To address this tension, we introduce Action QFormer, a query-based action-facing interface that uses instruction-conditioned queries to reorganize inherited multimodal information into action-facing representations before downstream action generation. In zero-shot sim-to-real navigation, Action QFormer improves average closed-loop task success from 18.8% to 56.3%, raises fixed-instruction action-generation correctness from 22.5% to 75.5%, and nearly eliminates out-of-distribution instruction generations. Further analyses show that Action QFormer changes how action supervision shapes inherited multimodal representations, reducing broad upstream rewriting while preserving targeted and sometimes constructive action-supervised adaptation. These results suggest that improving VLA performance requires not only stronger pretrained backbones, but also better ways of selecting and organizing inherited multimodal information while controlling how it is shaped under action supervision.