arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VersaCamVLA:用于机器人操作的相机可配置VLA策略

VersaCamVLA: Camera-Configurable VLA Policies for Robotic Manipulation

Boyao Han, Chen Shi, Jingjing Qian, ZhuoTan Tian, Li Jiang

arXiv 2610.12451首次发表:更新:

发表机构

The Chinese University of Hong Kong, Shenzhen; Harbin Institute of Technology, Shenzhen(香港中文大学(深圳); 哈尔滨工业大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLA模型部署时相机配置变化导致性能下降的问题,提出相机可配置框架VersaCamVLA,通过解耦相机集表示与动作学习实现稳健操作策略,在多基准及真实机器人上性能优于现有方法。

AI 中文摘要

视觉-语言-动作(Vision-Language-Action,VLA)模型已成为机器人操作领域强大的基础模型,但它们在训练期间依赖固定的相机配置,这使得它们在部署时对相机数量或位姿的变化十分脆弱。为克服这些局限,我们提出VersaCamVLA,一种相机可配置框架,它将相机集表示与动作学习解耦。VersaCamVLA学习统一的场景标记接口,该接口将任意、可变的带位姿RGB视图集映射为固定大小的潜在场景标记,这通过多信号目标视图预测和腕部增强位姿采样(Wrist-Augmented Pose Sampling,WAPS)实现,WAPS利用自然的腕部相机运动获取自由的位姿多样性。在部署时,轻量空间编码器将这些紧凑的场景标记作为补充视觉条件注入预训练的基础VLA,无需显式3D感知或新视图渲染。在RoboTwin、LIBERO和真实机器人平台上的实验表明,VersaCamVLA始终优于现有VLA方法和直接多视图基线,在不同相机数量和未见过的相机位姿下保持稳健性能。

英文摘要

Vision-Language-Action (VLA) models have emerged as powerful foundations for robotic manipulation, but their reliance on fixed camera configurations during training makes them brittle to changes in camera count or pose during deployment. To overcome these limitations, we propose VersaCamVLA, a camera-configurable framework that decouples camera-set representation from action learning. VersaCamVLA learns a unified scene-token interface that maps an arbitrary, variable set of posed RGB views into fixed-size latent scene tokens. This is achieved via multi-signal target-view prediction and Wrist-Augmented Pose Sampling (WAPS), which leverages natural wrist-camera motion for free pose diversity. At deployment, a lightweight spatial encoder injects these compact scene tokens into a pretrained base VLA as a supplementary visual condition, requiring no explicit 3D sensing or novel-view rendering. Experiments on RoboTwin, LIBERO, and a real-robot platform demonstrate that VersaCamVLA consistently outperforms prior VLA methods and direct multi-view baselines, maintaining robust performance across varying camera counts and unseen camera poses.

CommentsAccepted at NeurIPS 2026. Project page: https://boyaohan.github.io/VersaCamVLA.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑