arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34467cs.LG

对齐引导的流变换器用于高效视觉-语言-动作策略学习

Alignment-Guided Flow Transformer for Efficient Vision-Language-Action Policy Learning

  • Shanghai Jiao Tong University(上海交通大学)
  • Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Shengchao Hu, Peng Wang, Qiyang Zhou, Guodong Zheng, Yuqi Huang, Li Shen, Ya Zhang, Dacheng Tao

AI总结:

提出AGFT框架,通过显式三模态对齐损失和流匹配目标,解决VLA模型视觉-语言-动作错位问题,在基准上实现更高成功率和更低推理延迟。

AI中文摘要:

视觉-语言-动作(VLA)模型的最新进展通过统一感知、指令和控制,指向了通用机器人智能。尽管取得了令人瞩目的进展,现有的VLA模型往往因视觉、语言和动作之间的三模态错位而适应不良,这削弱了动作接地,损害了泛化能力和微调效率。在这项工作中,我们提出了对齐引导的流变换器(AGFT),一种新颖的框架,通过专门的对齐损失显式强制三模态对齐,弥合跨模态的表征差距并增强任务适应。虽然先前的研究主要强调双模态的视觉-语言对齐,我们系统地形式化并研究了VLA模型中的三模态对齐,并提供了消融实验和分析以隔离其在改善适应性和鲁棒性中的作用。为了进一步加速部署,我们采用流匹配目标,使得推理步骤显著少于基于扩散的策略,同时保持准确性。理论上,我们建立了三模态对齐差距与流匹配优化紧密度之间的定量联系;实证上,在广泛基准上的实验表明,AGFT相比最先进的基线实现了更高的成功率和更低的推理延迟,凸显了三模态对齐作为扩展鲁棒VLA操作的关键要素。

英文摘要:

Recent advances in Vision-Language-Action (VLA) models point toward general-purpose robotic intelligence by unifying perception, instruction, and control. Despite impressive progress, existing VLA models often adapt poorly due to \emph{tri-modal misalignment} among vision, language, and action, which weakens action grounding and hurts generalization and fine-tuning efficiency. In this work, we present Alignment-Guided Flow Transformer (AGFT), a novel framework that explicitly enforces tri-modal alignment through a dedicated alignment loss, bridging the representational gap across modalities and enhancing task adaptation. While prior research has predominantly emphasized bi-modal vision--language alignment, we systematically formalize and study tri-modal alignment in VLA models, and provide both ablations and analysis to isolate its role in improving adaptation and robustness. To further accelerate deployment, we adopt a flow-matching objective, enabling substantially fewer inference steps than diffusion-based policies while maintaining accuracy. Theoretically, we establish a quantitative connection between the tri-modal alignment gap and the optimization tightness of flow matching; empirically, experiments on the extensive benchmark show that AGFT achieves superior success rates and lower inference latency compared to SOTA baselines, underscoring tri-modal alignment as a key ingredient for scaling robust VLA manipulation.

补充信息

↑