arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向相机鲁棒性的视觉-语言-动作策略的跨视图动作一致性

Cross-View Action Consistency for Camera-Robust Vision-Language-Action Policies

Bingqi Huang, Bingchuan Wei, Xuan Wang, Yingkai Cai, Zhaokui Wang

arXiv 2608.06965首次发表:更新:

发表机构

Tsinghua University; Informatics Institute, University of Amsterdam; Faculty of Science, Vrije Universiteit Amsterdam(清华大学; 阿姆斯特丹大学信息学研究所; 阿姆斯特丹自由大学理学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对视觉-语言-动作策略的相机鲁棒性问题,提出跨视图动作一致性正则化方法,在LIBERO-Plus数据集及真实机器人任务上均显著提升了相机扰动下的策略性能。

AI 中文摘要

从固定场景摄像头微调得到的视觉-语言-动作(VLA)策略,在摄像头移动时可能失效,即便任务、物体、语言和机器人状态均未改变。本研究仅利用场景RGB图像、语言和本体感觉,不使用摄像头标签、外参、深度或点云输入,探究场景摄像头视角鲁棒性。全程屏蔽腕部视觉流,避免未受干扰的视觉捷径干扰对场景摄像头变化的归因。针对基于流的VLA策略,本文提出对动作-流速度场(直接积分生成连续动作块的量)进行正则化。通过将原始LIBERO演示重置为相同的MuJoCo状态,渲染标称与受扰动的场景摄像头视图,构建动作等价视图对。两种视图均受流匹配监督,同时跨视图损失鼓励其在相同采样流坐标处预测的动作-流速度一致。在LIBERO-Plus摄像头扰动赛道上,本方法达到87.2±0.4%的成功率(每个训练种子对应4797次rollout,共3个训练种子),较仅使用相同配对数据进行流匹配训练的基线(79.8±0.8%,同样3个种子)提升7.4个百分点,较朴素混合摄像头SFT提升12.5个百分点,同时保持标称摄像头ID性能(95.0±0.8%;相同数据的仅流匹配方法为95.0±4.3%)。打乱配对的对照实验结果降至25.8%,表明性能提升依赖于动作等价配对。在真实机器人上,针对3项桌面任务(每项任务及摄像头放置各10次rollout),在相同单场景RGB推理接口下,未见摄像头的成功率从53.3%提升至74.4%。

英文摘要

Vision-language-action (VLA) policies fine-tuned from a fixed scene camera can fail when the camera is moved, even when the task, objects, language, and robot state are unchanged. We study scene-camera viewpoint robustness using only a scene RGB image, language, and proprioception, without camera labels, extrinsics, depth, or point-cloud inputs. The wrist stream is masked throughout to prevent an unperturbed visual shortcut from confounding attribution to scene-camera variation. For flow-based VLAs, we propose to regularize the action-flow velocity field, the quantity directly integrated to generate continuous action chunks. We construct action-equivalent view pairs by resetting original LIBERO demonstrations to the same MuJoCo state and rendering nominal and perturbed scene-camera views. Both views are supervised by flow matching, while a cross-view loss encourages their predicted action-flow velocities to agree at the same sampled flow coordinates. On the LIBERO-Plus camera-perturbation track, our method reaches 87.2$\pm$0.4% (4,797 rollouts per seed across 3 training seeds), +7.4pp over flow-matching-only training on the same paired data (79.8$\pm$0.8%, also 3 seeds) and +12.5pp over naive mixed-camera SFT, while maintaining nominal-camera ID performance (95.0$\pm$0.8%; same-data FM-only: 95.0$\pm$4.3%). A shuffled-pair control collapses to 25.8%, showing that the gain depends on action-equivalent pairing. On a real robot, we evaluate three tabletop tasks with 10 rollouts per task and camera placement; held-out-camera success improves from 53.3% to 74.4% under the same single-scene-RGB inference interface.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑