发表机构
University of São Paulo; Instituto Israelita de Ensino e Pesquisa Albert Einstein; Graduate School of Information Sciences, Tohoku University(圣保罗大学; 以色列爱因斯坦教育与研究机构; 东北大学信息科学研究生院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对工业环境中机器人操作难题,提出共享自主性框架,利用RGB-D摄像头、视觉语言模型等,经GPU加速控制器等实现操作,在四足移动操纵器上验证,在多项任务中表现良好,展示了该框架的有效性。
AI 中文摘要
在工业环境中远程操作机器人操纵器需要高精度,仅基于摄像头的界面难以实现。操作员必须在有限深度感知下,在杂乱环境中将末端执行器与目标对齐,且不与周围结构碰撞。本文提出一个共享自主性框架来协助操作员。通过单个RGB-D摄像头捕捉操作员手臂动作和手势,无需可穿戴设备、基准或校准阶段。由视觉语言模型根据自由形式文本提示确定目标,并由可提示视频分割模型跟踪。GPU加速的模型预测控制器执行命令动作并避免碰撞,势场在最终接近时校正操作员参考。可通过手势触发自主模式。该框架在四足移动操纵器上得到验证,界面相对于运动捕捉地面真值的位置RMSE为59毫米,控制器使手臂与障碍物保持至少18厘米距离。在工业阀门操作和拾取放置任务中,完整框架在所有试验中成功,去除碰撞或协助模块会导致失败,自主执行在每个任务的五次试验中有四次成功。
英文摘要
Teleoperating a robotic manipulator in industrial environments demands precision that camera-based interfaces alone struggle to deliver. The operator must align the end-effector with a target in clutter, under limited depth perception, and without colliding with the surrounding structures. This paper presents a shared-autonomy framework that assists the operator throughout this process. A single RGB-D camera captures the operator's arm motion and hand gestures without wearables, fiducials, or a calibration stage. The intended target is specified by a free-form text prompt, grounded by a vision-language model in the robot's gripper camera, and tracked across its onboard cameras by a promptable video-segmentation model, resulting in a grasp frame continuously separated from the obstacle map. Every commanded motion is executed by a GPU-accelerated model-predictive controller that enforces self- and environment-collision avoidance against an online volumetric reconstruction, while a potential field corrects the operator's reference toward the grounded target during the final approach. An autonomous mode can be gesture-triggered to complete the grasp on the same target without a separate perception pipeline. The framework is validated on a quadruped mobile manipulator. The interface achieves a positional RMSE of 59 mm relative to motion-capture ground truth, and the controller keeps the arm at least 18 cm from obstacles while the operator deliberately commands the arm into them by 6 cm. In an industrial valve manipulation and a pick-and-place task, the full framework succeeded in all trials, while ablating either the collision or the assistance module produced failures through complementary mechanisms, and autonomous execution succeeded in four of five trials per task.