arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24180cs.RO

GraspTune:触觉驱动的执行细化实现稳健抓取

GraspTune: Tactile-Driven Execution Refinement for Robust Grasping

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • KTH Royal Institute of Technology(瑞典皇家理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Juntao Li, Xingke Xia, Sichao Liu, Daqiang Guo

中文总结 AI 辅助

GraspTune提出触觉驱动的执行细化框架,通过有界残余运动和学习接触语义,在模拟和真实机器人上显著提升多种视觉抓取候选生成器的稳定抓取成功率。

中文摘要 AI 辅助

视觉抓取候选生成已取得快速发展,但将选定的候选转化为稳定的物理抓取仍然是执行阶段的核心挑战。本文介绍了GraspTune,一种触觉驱动的执行阶段细化框架,该框架从名义候选出发,在接近、接触形成和最终抓取执行过程中应用有界的残余TCP运动。GraspTune利用状态条件专家接触查询和多任务监督(涵盖接触变化、接触风险和闭合后准备状态),从局部深度、触觉信号、状态和历史中学习面向控制的接触语义。该表示对扩散预训练的残差策略进行条件化,并与PPO对齐以实现闭环执行。在20个物体类别上超过60,000次模拟执行中,GraspTune在四种候选生成器上建立了执行层优势,将GraspNet、Contact-GraspNet、AnyGrasp和VGN的稳定抓取成功率分别提高了+19.22、+9.55、+12.45和+20.70个百分点。一项四倍留出类别研究将未见物体的执行成功率从54.58%提升至70.33%,展示了接触修正的类别不相交泛化能力。在配备Xense指尖传感器的UR5e平台上进行的超过1,000次真实机器人试验中,GraspTune将GraspNet的执行成功率从71.0%提升至84.3%,验证了无需真实世界策略微调的直接迁移能力。综合这些结果,该方法将视觉上合理的候选转化为稳定的物理抓取,以支持下游接触丰富的操作。补充视频可在以下https URL获取。

英文摘要

Visual grasp proposal generation has advanced rapidly, yet converting a selected proposal into a stable physical grasp remains a central execution-stage challenge. This paper introduces GraspTune, a tactile-driven execution-stage refinement framework that starts from a nominal proposal and applies bounded residual TCP motions during approach, contact formation, and final grasp execution. GraspTune learns control-facing contact semantics from local depth, tactile signals, state, and history using state-conditioned expert contact queries and multi-task supervision for contact change, contact risk, and post-close readiness. The representation conditions a diffusion-pretrained residual policy and is aligned with PPO for closed-loop execution. Across more than 60,000 simulated executions over 20 object categories, GraspTune establishes an execution-layer benefit across four proposal generators, raising stable grasp success by +19.22, +9.55, +12.45, and +20.70 percentage points for GraspNet, Contact-GraspNet, AnyGrasp, and VGN. A four-fold held-out category study raises unseen-object execution from 54.58% to 70.33%, showing category-disjoint generalization of contact correction. Across more than 1,000 real-robot trials on a UR5e setup with Xense fingertip sensors, GraspTune raises GraspNet execution from 71.0% to 84.3%, validating direct transfer without realworld policy fine-tuning. Together, these results turn visually plausible proposals into stable physical grasps for downstream contact-rich manipulation. A supplementary video is available at https://youtu.be/kcq7fSLNtzU.

↑