arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07553cs.RO

面向视觉机器人操纵的自动生成演示的自监督学习

Self Supervised Learning from Automatically Generated Demonstrations for Visual Robotic Manipulation

  • Technological University of Uruguay(乌拉圭技术大学)
  • Federal University of Santa Maria(圣玛丽亚联邦大学)
  • Federal University of Rio Grande(里奥格兰德联邦大学)

机构由 AI 辅助整理,请以论文原文为准。

Andres Rivas, Anselmo R. Cukla, Rodrigo S. Guerra, Bruna V. Guterres, Ricardo B. Grando

AI总结:

该研究提出一种基于ROS 2和Isaac Sim的自监督视觉机器人操纵方法,通过自动生成演示训练卷积网络,在仿真和真实UR5e机器人上验证了其抓取成功率,为视觉操纵提供了低部署成本的方案。

AI中文摘要:

机器人操纵通常需要针对特定对象的编程、手动数据标注或校准的感知流水线,这限制了其在实际场景中的快速部署。从演示中学习提供了一种更直接的替代方案,但收集演示仍可能需要人类远程操作或运动学示教。本文提出一种自监督视觉操纵方法,其中机器人围绕目标位姿自动生成演示,并直接从腕部安装的RGB图像学习相对位姿修正。所提出的流水线使用ROS 2和Isaac Sim收集带标签的图像-位姿对,无需明确的相机到机器人的外部校准。为平面微调与粗略三维接近分别生成数据集,并训练卷积网络从单帧RGB观测中回归相对平移与旋转。执行期间,由粗到精控制器先使用经高度变化训练的模型接近对象,再用平面数据优化最终对齐。该方法在仿真环境和配备夹爪与单目相机的真实UR5e协作机器人上均进行了评估。在仿真中,微调阶段将最终平面离散度从9.69 mm降至5.38 mm;在真实实验中,系统对三个物理对象执行端到端抓取尝试,其中两个无旋转对象的成功率达66.6%和63.6%,在旋转条件下仍保持部分鲁棒性。这些结果表明,自动生成的演示可在有限设置成本下支持实用的视觉操纵,同时也揭示了深度预测和对象依赖泛化方面的剩余挑战。

英文摘要:

Robotic manipulation often requires object specific programming, manual data annotation, or calibrated perception pipelines, which limits rapid deployment in practical settings. Learning from demonstration offers a more direct alternative, but collecting demonstrations can still demand human teleoperation or kinesthetic teaching. This paper presents a self supervised visual manipulation method in which a robot automatically generates demonstrations around a target pose and learns relative pose corrections directly from wrist mounted RGB images. The proposed pipeline uses ROS~2 and Isaac Sim to collect labeled image-pose pairs without requiring explicit camera to robot extrinsic calibration. Separate datasets are generated for planar refinement and coarse three dimensional approach, and a convolutional network is trained to regress relative translation and rotation from single frame RGB observations. During execution, a coarse to fine controller first approaches the object using models trained with height variation and then refines the final alignment using planar data. The method is evaluated both in simulation and on a real UR5e collaborative robot equipped with a gripper and a monocular camera. In simulation, the refinement stage reduces the final planar dispersion from 9.69 mm to 5.38 mm. In real world experiments, the system performs end to end grasp attempts on three physical objects and reaches success rates of 66.6% and 63.6% for two objects without object rotation, while still maintaining partial robustness under rotated conditions. These results show that automatically generated demonstrations can support practical visual manipulation with limited setup effort, while also exposing remaining challenges in depth prediction and object dependent generalization.

补充信息

↑