arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAGP:基于粗区域VLM推理的语义 affordance 引导抓取规划

SAGP: Semantic Affordance-Guided Grasp Planning via Coarse-Zone VLM Reasoning

Muhayy Ud Din, Irfan Hussain

arXiv 2607.29374首次发表:更新:

发表机构

Khalifa University; KU Center for Autonomous Robotic Systems (KUCARS)(哈利法大学; KU自主机器人系统中心(KUCARS))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SAGP是一种无训练的抓取规划框架,通过粗区域抽象层结合VLM推理与几何规划,在保持纯几何抓取高成功率的同时,提升了抓取的功能适用性,尤其适用于非对称带把手物体。

AI 中文摘要

基于几何的抓取规划器可保证抓取的物理有效性,但忽略功能语义,常生成对跖且无碰撞却不实用的抓取,比如抓马克杯杯口、刀刃、瓶盖附近的瓶子,即便满足传统抓取指标,也会导致下游任务失败。现有视觉语言模型(VLM)方法要么依赖细粒度、类别特定的部件分割,要么尝试直接推断抓取位姿,后者易出现空间幻觉。目前尚无实用的无训练框架能将高级语义推理与几何抓取规划稳健关联。我们提出语义 affordance 引导抓取规划(SAGP),这是一种基于粗区域抽象层的无训练流水线。该方法先通过基于PCA的对齐,再用距离驱动的DBSCAN聚类,将物体点云划分为空间区域(顶部、中部、底部、侧面和凸起),完全绕过学习到的分割。预训练VLM随后通过结构化零样本查询评估每个区域的抓取质量,将得到的区域分数与几何、可达性及任务对齐信号融合,对跖抓取候选进行重新排序。在PyBullet中使用Franka Panda机器人对YCB物体开展的实验显示,SAGP保留了纯几何规划的高成功率,同时大幅提升了所选抓取的功能适用性,尤其在纯几何信息不足的非对称带把手物体上表现突出。引入的粗区域抽象为基于VLM的推理与几何抓取规划提供了有效的无训练桥梁,无需细粒度部件分割。

英文摘要

Geometry-based grasp planners ensure physically valid grasps but ignore functional semantics, often generating grasps that are antipodal and collision-free yet practically inappropriate, for example, gripping a mug by its rim, a knife by the blade, or a bottle near its cap. These inconsistencies cause the downstream task to fail even when traditional grasp metrics are met. Existing vision-language model (VLM) approaches either depend on fine-grained, category-specific part segmentation or attempt to directly infer grasp poses, with the latter prone to spatial hallucinations. As a result, no practical, training-free framework has yet been proposed that robustly links high-level semantic reasoning to geometric grasp planning. We introduce Semantic Affordance-Guided Grasp Planning (SAGP), a training-free pipeline built on a coarse-zone abstraction layer. The method first partitions the object point cloud into spatial regions (top, middle, bottom, lateral sides, and protrusions) by applying PCA-based alignment followed by distance-driven DBSCAN clustering, entirely bypassing learned segmentation. A pre-trained VLM then assesses the grasp quality of each region through a structured zero-shot query, and the resulting zone-wise scores are fused with geometric, reachability, and task-alignment signals to re-rank antipodal grasp candidates. Experiments on YCB objects in PyBullet with a Franka Panda robot show that SAGP preserves the high success rate of geometry-only planning while substantially improving the functional appropriateness of selected grasps, particularly on asymmetric, handle-bearing objects where geometry alone is uninformative. The introduced coarse-zone abstraction offers an effective, training-free bridge between VLM-based reasoning and geometric grasp planning, without the need for fine-grained part segmentation.

CommentsAccepted in ICAME2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑