Fisheye-VLA:利用单个鱼眼相机解耦覆盖范围与视觉敏锐度以实现操作
Fisheye-VLA: Decoupling Coverage and Acuity for Manipulation with a Single Fisheye Camera
浏览论文内容
中文总结 AI 辅助
本文提出Fisheye-VLA,利用单个鱼眼相机结合全局视图与末端执行器为中心的局部裁剪,在扩展桌面区域实现84%和82%的操作成功率,无需腕部相机。
中文摘要 AI 辅助
操作任务既需要广泛的场景感知,也需要详细的局部反馈,然而传统的相机装置通过分开的前置相机和腕部相机来提供这两种能力。我们提出了Fisheye-VLA,一种视觉接口,它利用单个被动式鱼眼相机将这些能力结合在一起。全局视图保留了工作空间,而局部透视裁剪则将细节导向交互区域。关键的设计问题在于这个局部视觉预算应该投向何处。我们通过一项受控的重新渲染研究来回答这个问题,在相同的记录观测数据上比较了不同的裁剪方向。研究发现,以末端执行器为中心的视图捕获了更大候选池中估计收益的大部分,这促使我们在双手周围进行紧凑的分配。我们的接口使用校准的末端执行器投影和运动引导来跟踪裁剪区域,同时共享的光线编码在裁剪区域移动时保持其空间意义。与预训练的VLA集成后,在两个扩展的桌面区域中实现了84%和82%的成功率,其中一些目标放置超出了前置相机的覆盖范围,并支持货架和传送带操作。消融研究表明,在更大的工作空间区域中,局部裁剪及其观看方向变得更加重要。结果表明,单个鱼眼相机可以在没有物理腕部相机的情况下支持这些操作任务。
英文摘要
Manipulation requires both broad scene awareness and detailed local feedback, yet conventional camera rigs provide them through separate front and wrist cameras. We present Fisheye-VLA, a visual interface that brings these capabilities together using a single passive fisheye. A global view preserves the workspace, while local perspective crops direct detail toward the interaction. The key design question is where this local visual budget should go. We answer it through a controlled re-rendering study, comparing alternative crop directions on the same recorded observations. The study finds that end-effector-centered views capture most of the estimated benefit of a much larger candidate pool, motivating a compact allocation around both hands. Our interface uses calibrated end-effector projection and motion lead to track the crops, while a shared ray encoding preserves their spatial meaning as they move. Integrated with a pretrained VLA, it achieves 84% and 82% success in the two expanded tabletop regions, where some target placements extend beyond the front-camera coverage, and supports shelf and conveyor manipulation. Ablations show that local crops and their viewing directions become more important in the larger workspace regions. The results demonstrate that a single fisheye can support these manipulation tasks without physical wrist cameras.
发表机构
- Hong Kong Embodied AI Lab(香港具身智能实验室)
- The Chinese University of Hong Kong(香港中文大学)
- DeepCybo
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。