InfraVLA:利用基础设施摄像头扩展视觉-语言-动作导航
InfraVLA: Extending Vision-Language-Action Navigation with Infrastructure Cameras
- University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
InfraVLA通过引入CCTV编码器将预训练VLA导航模型适配基础设施摄像头,两阶段训练提升决策,在模拟和真实任务中显著优于基线。
中文摘要 AI 辅助
许多机器人运行的室内环境(如仓库、办公室和医院)已经安装了摄像头。这些摄像头能够观察到机器人所处位置无法看到的建筑部分,然而导航策略(包括最新的视觉-语言-动作(VLA)模型)并未利用这些摄像头。我们提出InfraVLA,一种端到端方法,将预训练的导航VLA适配到这类静态基础设施视图:一个闭路电视(CCTV)编码器将每个外部视图转换为输入序列的令牌。由于这些视图仅在少数决策点起作用,仅靠微调在我们的实验中并未使策略利用它们;因此我们分两个阶段训练,先使用带有上采样反事实数据的演示,再使用恢复数据。我们在两个模拟仓库任务上评估:寻找指令中命名的物体和绕行被堵塞的通道,其中决定性信息往往仅对基础设施摄像头可见。在分布内测试中,InfraVLA在两个任务上均达到100%的成功率,而基线(无CCTV输入)分别为34.0%和73.6%。在分布外测试集上,它分别达到88.2%和88.9%。在真实四足机器人上,使用不到10分钟的演示进行微调后,策略达到83.3%,而仅使用机载传感器的基线为29.2%。
英文摘要
Many indoor environments in which robots operate, such as warehouses, offices, and hospitals, already have cameras installed. They observe parts of the building that the robot cannot see from where it stands, yet navigation policies, including recent vision-language-action (VLA) models, do not use them. We propose InfraVLA, an end-to-end method that adapts a pretrained navigation VLA to such static infrastructure views: a closed-circuit television (CCTV) encoder turns each external view into tokens of the input sequence. Because the views matter only at rare decision points, fine-tuning alone did not make the policy use them in our experiments; we therefore train in two stages, on demonstrations with upsampled counterfactual data and then on recovery data. We evaluate on two simulated warehouse tasks, finding an object named in the instruction and rerouting around blocked aisles, where the deciding information is often visible only to the infrastructure cameras. Tested in distribution, InfraVLA reached a success rate of 100% on both, against 34.0% and 73.6% for a baseline without CCTV input. On out-of-distribution test sets it reached 88.2% and 88.9%. On a real quadruped fine-tuned with under 10 minutes of demonstrations, the policy reached 83.3% against 29.2% for the on-board-only baseline.