arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于可组合基础先验与通用抓取合成的自适应视觉-语言抓取

Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

Sixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang, Guohua Fan, Fan Huang, Haoxuan Li, Yongkang Li, Yuhan Li, Bencheng Liao, Zeyu Zhang, Wenyu Liu, Hangxin Liu, Xinggang Wang

arXiv 2609.04096首次发表:更新:

发表机构

School of Electronic Information and Communications, Huazhong University of Science and Technology; KEENON Robotics Co., Ltd.; Suzhou Silicon Era Intelligent Technology Co., Ltd.; College of Materials, Xiamen University; School of Artificial Intelligence and Automation, Huazhong University of Science and Technology; ByteDance; State Key Laboratory of General Artificial Intelligence, Beijing Institute for General Artificial Intelligence (BIGAI); Hubei Automation Institute(华中科技大学电子信息与通信学院; 科语机器人有限公司; 苏州硅时代智能科技有限公司; 厦门大学材料学院; 华中科技大学人工智能与自动化学院; 字节跳动; 北京通用人工智能研究院通用人工智能国家重点实验室; 湖北省自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出AdaRoboVLG框架,通过解耦物理抓取合成与任务依赖理解,结合可组合基础先验实现自适应通用抓取,在仿真与真实实验中展现出高效学习、跨机械臂泛化及应对复杂环境的能力。

AI 中文摘要

本文提出AdaRoboVLG,这是一种支持不同机械臂通用抓取合成的任务自适应视觉-语言-抓取(VLG)框架。与现有将基础模型与端到端抓取策略紧密耦合的VLG方法不同,AdaRoboVLG学习一种高效的通用基础策略,该策略通过显式运动学映射和基于力闭合的稳定性估计生成并评估物理可行的抓取候选,同时将依赖任务的理解任务交由专门的基础模型模块处理。这些模块提供可组合的先验,这些先验被集成到抓取合成过程中,从而在不重新训练底层抓取策略的情况下实现上下文自适应的抓取合成。通过大量仿真和真实世界实验,我们证明:(i)该基础策略展现出高效的学习能力和强大的跨机械臂泛化能力;(ii)该框架能有效整合空间、认知和时间先验,以应对三种代表性抓取挑战,且与最先进方法相比,抓取合成性能未受影响;(iii)这些先验可协同运作,以在杂乱和动态环境中实现功能性抓取。这些结果表明,将物理抓取合成与依赖任务的理解解耦,为机器人抓取提供了一种可扩展的范式,使得基础模型的未来进展可直接转化为抓取能力的提升,而无需重新设计或重新训练底层抓取策略。补充视频可在该https URL获取。

英文摘要

This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient generalizable base policy that generates and evaluates physically feasible grasp candidates through explicit kinematic mapping and force-closure-based stability estimation, while offloading task-dependent understanding to specialized foundation-model modules. These modules provide composable priors that are integrated into the grasp synthesis process, enabling contextually adaptive grasp synthesis without retraining the underlying grasp policy. Through extensive simulation and real-world experiments, we demonstrate that (i) the base policy exhibits efficient learning and strong cross-hand generalization, (ii) the framework effectively incorporates spatial, cognitive, and temporal priors to address three representative grasping challenges without compromising grasp synthesis performance compared to state-of-the-art methods, and (iii) these priors can operate jointly to enable functional grasping in cluttered and dynamic environments. These results indicate that decoupling physical grasp synthesis from task-dependent understanding provides a scalable paradigm for robotic grasping, allowing future advances in foundation models to be directly translated into improved grasp capabilities without redesigning or retraining the underlying grasp policy. Supplementary videos are available at https://adarobovlg.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑