多视图统一相机场:面向仅RGB多相机视觉-语言-动作(VLA)策略的几何形状动作表征
Multi-View Unified Camera Fields: Geometry-Shaped Action-Facing Representations for RGB-Only Multi-Camera VLA Policies
浏览论文内容
中文总结 AI 辅助
本研究针对多相机VLA策略的动作表征缺陷,提出仅训练用的MVUCF框架,通过几何相关目标优化后部署无额外推理开销,在LIBERO、RoboTwin等任务及真实人形实验中均取得显著性能提升。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型在机器人操纵中展现出强大的泛化能力,但复杂的接触密集型任务通常需要多相机观测,以联合捕捉遮挡下的末端执行器、物体和目标。现有多相机VLA通常拼接视图token,导致动作表征在度量深度上较弱且跨相机不一致。我们提出多视图统一相机场(MVUCF),这是一个仅用于训练的框架,可在各视图间形成面向动作的共享潜在场。坐标查询深度目标使度量深度可恢复,而预处理感知对应目标可对齐不同相机观测同一物理点的token,两者直接塑造动作模块使用的隐藏状态。注入几何信息后,深度、相机标定和辅助头被移除,因此部署时使用原始仅RGB图,无额外推理浮点运算量(FLOPs)。保留的探针证实其深度恢复和跨视图匹配能力更强。在匹配的GR00T-N1.6设置下,MVUCF在LIBERO上达到98.9%,使LIBERO-Plus提升22.4个百分点,并在涵盖触摸、移动放置和接触交互三个动作类别的六项RoboTwin任务中,成功率提升23.3个百分点。真实世界人形机器人实验进一步证明了其在仅RGB部署下的实际有效性。
英文摘要
Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation, yet complex contact-rich tasks often benefit from multi-camera observations that jointly capture the end effector, objects, and targets under occlusion. Existing multi-camera VLAs usually concatenate view tokens, leaving action representations weak in metric depth and inconsistent across cameras. We introduce Multi-View Unified Camera Fields (MVUCF), a training-only framework that forms a shared action-facing latent field across views. A coordinate-query depth objective makes metric depth recoverable, while a preprocessing-aware correspondence objective aligns tokens observing the same physical point from different cameras. Both directly shape the hidden states consumed by the action module. After geometry injection, depth, camera calibration, and auxiliary heads are removed, so deployment uses the original RGB-only graph with no extra inference FLOPs. Held-out probes confirm stronger depth recovery and cross-view matching. Under matched GR00T-N1.6 settings, MVUCF reaches 98.9% on LIBERO, improves LIBERO-Plus by 22.4 points, and raises success by 23.3 points across six RoboTwin tasks spanning three action families: touch, move-and-place, and contact interaction. Real-world humanoid experiments further provide evidence of its practical effectiveness under RGB-only deployment.