Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis
基于可组合基础先验与通用抓取合成的自适应视觉-语言抓取
机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院) ; KEENON Robotics Co., Ltd.(科语机器人有限公司) ; Suzhou Silicon Era Intelligent Technology Co., Ltd.(苏州硅时代智能科技有限公司) ; College of Materials, Xiamen University(厦门大学材料学院) ; School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) ; ByteDance(字节跳动) ; State Key Laboratory of General Artificial Intelligence, Beijing Institute for General Artificial Intelligence (BIGAI)(北京通用人工智能研究院通用人工智能国家重点实验室) ; Hubei Automation Institute(湖北省自动化研究所)
AI总结 本文提出AdaRoboVLG框架,通过解耦物理抓取合成与任务依赖理解,结合可组合基础先验实现自适应通用抓取,在仿真与真实实验中展现出高效学习、跨机械臂泛化及应对复杂环境的能力。