SE(3) 神经势场:无需显式3D重建,直接从图像进行6自由度轨迹规划
SE(3) Neural Potential Fields for 6-DoF Trajectory Planning Directly from Images Without Explicit 3D Reconstruction
- University of Maine(缅因大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出从RGB图像学习、以测地距离监督的SE(3)神经势场,无需3D重建即可规划6自由度无碰撞轨迹,显著提升成功率与规划速度。
AI中文摘要:
在杂乱环境中达到6自由度抓取位姿需要无碰撞轨迹,传统方法通过3D重建场景并在重建结果内进行规划获得,但代价是重建的精度和计算开销。直接从图像学习的势场消除了这种依赖,但继承了人工势场的经典缺陷:当吸引梯度和排斥梯度相互抵消时,下降轨迹会擦过障碍物而非绕行,并可能在到达目标前停滞。我们提出了一种从带位姿的RGB图像学习、并以导航函数(即训练时从相同图像恢复的自由空间中到抓取点的测地距离)监督的SE(3)神经势场,该势场消除了上述两种失败。在两个桌面场景中,从障碍物阻挡的起始点开始,在UR10上执行,该势场从每个起始点都收敛到距抓取点3厘米以内,且其执行的每条路径相对于真实几何都是无碰撞的,而在仅图像监督下,这两个比例分别为25%和0%;平均间隙从不足1厘米提高到8.6-8.8厘米,机械臂连杆接触从执行配置的20.6%-50.4%降至2.7%-5.5%。在两个场景中,执行的抓取成功率分别为90.0%和40.0%,残余失败是笛卡尔执行器的拒绝而非势场的问题。规划耗时约2秒,而RRT*在相同图像的重建上需67-133秒,尽管在常见的离线测试框架下两者相当:部署时的差距来自密集重建的碰撞检测成本,而非规划器的复杂度。
英文摘要:
Reaching a 6-DoF grasp pose in clutter requires a collision-free trajectory, conventionally obtained by reconstructing the scene in 3D and planning inside that reconstruction, at the cost of its accuracy and compute. Potential fields learned directly from images remove that dependency but inherit the classical weakness of artificial potential fields: where attractive and repulsive gradients cancel, the descent grazes the obstacle instead of going around it, and can stall short of the goal. We present an SE(3) neural potential field learned from posed RGB images and supervised with a navigation function, the geodesic distance to the grasp through free space recovered from those same images during training, which removes both failures. On two tabletop scenes, from obstacle-blocked starts executed on a UR10, the field converges within 3 cm of the grasp from every start and every path it executes is collision-free against the ground-truth geometry, against 25% and 0% under image supervision alone; mean clearance rises from under a centimeter to 8.6-8.8 cm and arm-link contacts fall from 20.6-50.4% to 2.7-5.5% of executed configurations. Executed grasp success is 90.0% and 40.0% on the two scenes, the residual failures being refusals of the Cartesian executor rather than of the field. Planning takes about 2 s against 67-133 s for RRT* on a reconstruction of the same images, though under a common offline harness the two are comparable: the deployed margin is the cost of collision-checking a dense reconstruction, not planner complexity.