arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WARP-VLA:面向视觉-语言-动作模型中视角鲁棒策略执行的腕部相机适配

WARP-VLA: Wrist-Camera Adaptation for View-Robust Policy Execution in Vision-Language-Action Models

Junmyeong Lee, Dongmin Shin, Min-Gyu Park, Wooseok Jeon, Inho Chang, Hae-Gon Jeon

arXiv 2610.11508首次发表:更新:

发表机构

Yonsei University; Korea Electronics Technology Institute(KETI)(延世大学; 韩国电子技术研究院(KETI))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉-语言-动作模型在机器人操作中对相机配置变化敏感的问题,提出WARP-VLA,采用混合专家架构实现腕部视角鲁棒策略,在LIBERO基准上将pi-0.5平均成功率从39.2%提至78.3%,且可迁移至真实机器人。

AI 中文摘要

尽管近期面向机器人操作的视觉-语言-动作模型(VLAs)取得了进展,但其性能仍对相机配置的变化敏感,在跨设置部署时问题更为明显,因为几乎不可能复现训练所用的精确相机位姿。与固定外部视角不同,腕部视角更具挑战性,因为相机随机器人移动,即使微小的安装变化也会改变细粒度几何线索。为解决该问题,我们提出WARP-VLA,一种面向多种腕部相机配置的相机视角鲁棒VLA。WARP-VLA采用混合专家(MoE)架构,其中各专家学习视角特定的特征变换,路由模块基于隐式视角信息组合这些专家,使策略部署无需将相机外参作为额外输入。通过在LIBERO基准上的实验,WARP-VLA在腕部视角扰动下将pi-0.5的平均成功率从39.2%提升至78.3%。真实机器人实验进一步表明,在仿真中学习到的特征级适配可成功迁移至多种部署设置。为便于可复现性和未来研究,我们发布了腕部视角鲁棒性基准及即插即用实现。

英文摘要

Despite recent advances in Vision-Language-Action models (VLAs) for robotic manipulation, their performance remains sensitive to changes in camera configuration. The problem becomes more evident in cross-setup deployment, as reproducing the exact camera pose used for training is nearly impossible. Unlike fixed external views, wrist views are more challenging because the camera moves with the robot, causing even small mounting variations to alter fine-grained geometric cues. To address this, we propose WARP-VLA, a camera-view robust VLA for diverse wrist camera configurations. WARP-VLA adopts a Mixture-of-Experts (MoE) architecture where individual experts learn view-specific feature transformations, and a router combines them based on implicit view information. This allows the policy to be deployed without requiring camera extrinsic parameters as additional input. Through experiments on the LIBERO benchmark, WARP-VLA improves the average success rate of pi-0.5 from 39.2% to 78.3% under wrist-view perturbations. The real-robot experiments further show that the feature-level adaptation learned in simulation successfully transfers to diverse deployment settings. To facilitate reproducibility and future research, we release our wrist viewpoint robustness benchmark and a plug-and-play implementation.

Comments9 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑