arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14341cs.ROcs.AI

超越视觉抓取:从检测到执行的复杂抓取基准测试

Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution

Hanyi Zhang, Khang Nguyen, Charith Munasinghe, Basu Hela, Tianyu Li, Zihong Luo, Hoan Nguyen, Hans Wernher van de Venn, Yalin Zheng, Ravi Prakash, Tung D. Ta, A… 展开作者

Hanyi Zhang, Khang Nguyen, Charith Munasinghe, Basu Hela, Tianyu Li, Zihong Luo, Hoan Nguyen, Hans Wernher van de Venn, Yalin Zheng, Ravi Prakash, Tung D. Ta, Anh Nguyen, Baoru Huang

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有抓取基准未涵盖复杂多步推理和语义理解任务的问题,提出GCA-Bench基准,实现多种基线方法并进行实证研究,给出新评估指标等,为开发更强大通用的抓取策略提供指导。

中文摘要 AI 辅助

强大的机器人抓取对于复杂的现实世界应用仍然是一个基本挑战。大规模模型的最新进展展示了在机器人任务中推理的潜力。然而,现有的抓取基准主要集中在孤立的基于视觉的抓取姿态检测,未能涵盖执行过程中需要多步推理和语义理解的抓取任务的复杂性。为填补这一空白,我们提出了GCA-Bench,这是一个具有挑战性的“复杂动作抓取”场景的基准,涉及场景级推理和语义约束。GCA-Bench能在相同设置下评估近期大型基础模型。为证明新基准的有效性,我们实现了从传统抓取检测管道到端到端学习方法的各种基线。实证研究在复杂抓取场景中的成功率低于70%,凸显了关键局限性。此外,我们提出了新的评估指标,分析了关键失败模型,并提供了见解以指导更强大和通用的抓取策略的发展。

英文摘要

Robust robotic grasping remains a fundamental challenge for complex real-world applications. Recent advances in large-scale models demonstrate promising capabilities for reasoning in robotic tasks. However, existing benchmarks for grasping primarily focus on isolated, visual-based grasp pose detection, failing to capture the complexity of grasping tasks that require multi-step reasoning and semantic understanding during execution. To address this gap, we propose GCA-Bench, a benchmark featuring challenging \textit{grasping with complex action} scenarios that involve both scene-level reasoning and semantic constraints. GCA-Bench enables the evaluation of recent large foundation models under the same settings. To demonstrate the effectiveness of our new benchmark, we implement a diverse set of baselines, ranging from traditional grasp detection pipelines to end-to-end learning methods. Empirical studies achieve success rates below 70\% on complex grasping scenarios, underscoring critical limitations. In addition, we propose new evaluation metrics, analyze critical failure models, and provide insights to guide the development of more robust and generalizable grasping strategies.

发表机构

  • University of Liverpool(利物浦大学)
  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
  • Zurich University of Applied Science(苏黎世应用科学大学)
  • Indian Institute of Science(印度科学研究所)
  • University of Information Technology(信息技术大学)
  • The University of Tokyo(东京大学)

机构由 AI 辅助整理,请以论文原文为准。

↑