LAST: Bridging Vision-Language and Action Manifolds via Gromov-Wasserstein Alignment
LAST: 通过Gromov-Wasserstein对齐连接视觉-语言与动作流形
机构 * National Engineering Laboratory for Intelligent Information Processing(智能信息处理国家级工程实验室) ; University of Science and Technology of China(中国科学技术大学)
专题命中 VLA模型 :VLA(summary_cn,abstract);vision-language-action(abstract);分类 cs.CV
AI总结 提出LAST方法,通过李代数线性化和局部度量离散化,对齐视觉-语言语义几何与动作流形,解决异构空间不兼容问题,提升VLA模型收敛性和泛化性。