发表机构
University of California, Merced; University of California, Davis(加州大学默塞德分校; 加州大学戴维斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究首次提出多臂机器人水果采摘的视觉语言模型基准测试,验证了VLM零样本规划的有效性,并指出3D路径点生成与碰撞协调是实际部署的主要瓶颈。
AI 中文摘要
多臂机器人采摘为提高采摘效率和减少对人工劳动的依赖提供了一条有前景的途径。然而,实际部署仍然具有挑战性,因为系统必须在不同环境中进行泛化,同时在共享工作空间中高效协调多个机械臂。现有方法通常需要在目标环境中进行大量数据收集,或依赖简化假设,从而限制了规划质量。在这项工作中,我们首次引入了一个全面的基准测试,用于评估预训练的视觉语言模型(VLMs)在多臂水果采摘规划中的零样本(zero-shot)表现。我们的基准测试使用真实的苹果和柑橘果园图像,并将基于VLM的规划流程与传统感知与规划流程进行比较。VLM流程直接为每个机械臂生成采摘序列和路径点,而轻量级轨迹验证器则检查碰撞情况。我们的结果表明,前沿VLM能够零样本生成有效的多臂采摘计划,但实际部署仍受限于精确的3D路径点生成和碰撞感知协调。这些结果凸显了预训练VLM在多臂机器人采摘中的潜力与当前局限性。
英文摘要
Multi-arm robotic harvesting offers a promising path to improve harvesting efficiency and reduce reliance on manual labor. However, practical deployment remains challenging because the system must generalize across diverse environments while efficiently coordinating multiple arms in a shared workspace. Existing methods often require substantial data collection in target environments or rely on simplifying assumptions that limit planning quality. In this work, we introduce the first comprehensive benchmark for evaluating pretrained Vision-Language Models (VLMs) on zero-shot multi-arm fruit harvesting planning. Our benchmark uses real-world apple and citrus orchard images and compares a VLM-based planning pipeline with a traditional perception-and-planning pipeline. The VLM pipeline directly generates harvesting sequences and waypoints for each arm, while a lightweight trajectory verifier checks for collisions. Our results show that frontier VLMs can generate effective multi-arm harvesting plans zero-shot, but a practical deployment remains limited by accurate 3D waypoint generation and collision-aware coordination. These results highlight both the promise and current limitations of pretrained VLMs for multi-arm robotic harvesting.