发表机构
Gachon University(嘉泉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GeoBridge-VLA通过两阶段方法从冻结的视觉编码器中学习几何特征,利用门控残差增强视觉标记,无需深度输入,在LIBERO和实体机器人上均显著优于SmolVLA。
AI 中文摘要
视觉-语言-动作(VLA)模型编码了来自视觉-语言预训练中的语义信息,但操作任务还需要精确的空间推理。我们提出了GeoBridge-VLA,一种两阶段方法,用于从预训练VLA的冻结视觉编码器中学习几何特征,并将其用于动作预测。第一阶段训练一个特征桥接模块和几何解码器,使用深度监督。第二阶段冻结这些模块,并训练一个门控残差接口以及动作侧的投影和动作专家。该残差增强了现有的视觉标记,而不增加第二个图像编码器或增加标记数量。部署需要RGB图像、机器人状态和语言,但不需要深度观测。在匹配的评估条件下,GeoBridge-VLA在LIBERO基准上达到了70.9%的成功率,而SmolVLA为60.0%。在同一训练检查点中禁用残差,成功率从70.90%降至69.85%,且在不同套件中效果不一。在物理ROBOTIS OMY机器人上,GeoBridge-VLA在四个任务中的200次试验中成功了148次(74.0%),而SmolVLA为200次中的108次(54.0%)。
英文摘要
Vision-language-action (VLA) models encode semantic information from vision-language pretraining, but manipulation also requires precise spatial reasoning. We present GeoBridge-VLA, a two-stage method for learning geometric features from a pretrained VLA's frozen visual encoder and using them for action prediction. Stage I trains a feature bridge and geometry decoder with depth supervision. Stage II freezes these modules and trains a gated residual interface together with the action-side projections and action expert. The residual augments the existing visual tokens without adding a second image encoder or increasing the token count. Deployment requires RGB, robot state, and language, but no depth observations. Under matched evaluation conditions, GeoBridge-VLA achieves 70.9% success on LIBERO, compared with 60.0% for SmolVLA. Disabling the residual in the same trained checkpoint reduces success from 70.90% to 69.85%, with mixed effects across suites. On a physical ROBOTIS OMY robot, GeoBridge-VLA succeeds in 148 of 200 trials (74.0%) across four tasks, compared with 108 of 200 (54.0%) for SmolVLA.