发表机构
University of Nottingham(诺丁汉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出螺旋注意力层,以空间变换表示物体间关系,实现刚体力学等变消息传递,在 LIBERO-Spatial 上以极少参数超越大网络,并提升控制器性能。
AI 中文摘要
学习型操作策略从数据中重新发现刚体力学以闭式形式提供的空间关系。这耗费数据,并使策略对场景中的几何变化变得脆弱。我们提出螺旋注意力(Screw Attention),一种 transformer 层,其中两个物体之间的关系是空间变换而非图边。每个 token 都是一个带有位姿的物体。每对 token 携带相对位姿,对于机器人关节,还携带关节螺旋。消息沿此关系传输到接收者的坐标系,而注意力分数仅看到坐标系不变的数量。通过构造,消息对每个 token 处的独立坐标系变化是等变的,单层即可表达刚体力学的速度递归。在模拟操作任务中,螺旋注意力在 LIBERO-Spatial 上匹配或超越同规模的控制网络,包括图网络、transformer 网络和扁平网络。凭借 16,162 个参数,它从物体位姿(无图像或语言)达到 97.3% 的成功率,高于参数多 27 倍的扁平网络。在每连杆坐标系约定变化下,其成功率不变,而所有其他学习网络降至 3% 以下。作为门控残差置于解析控制器上,它将插入成功率提高 17.3 个百分点。它不受高达 10 毫米的位姿噪声和 Franka 机械臂出厂校准范围内的关节偏移的影响。这些结果提出一个标准:当任务需要系统其他部分未提供的坐标系间关系时,几何是决定性的。代码和训练好的策略将发布。
英文摘要
Learned manipulation policies rediscover from data the spatial relations that rigid-body mechanics supplies in closed form, which leaves them fragile to geometric change. We present Screw Attention, a transformer layer in which the relation between two bodies is a spatial transform rather than a graph edge. Each pair of tokens carries the relative pose and, for robot joints, the joint screw. Messages are transported along this relation into the receiver's frame, while the attention scores see only frame-invariant quantities. By construction, the messages are equivariant to an independent change of frame at every token, and a single layer can express the velocity recursion of rigid-body mechanics. On LIBERO-Spatial, a policy of 16k parameters trained from object poses alone reaches 97.3% success, above graph, transformer and flat networks of the same size and a flat network with 27 times more parameters. Ablations show that the gain comes from transporting the correct relations, and that the structure pays most where the task requires relations between frames that nothing else supplies. The equivariance makes the policy robust to how the robot is described, where every other learned network collapses under a change of frame convention. Furthermore, the policy tolerates pose noise and calibration errors at least as well as an analytic controller. Used as a gated residual on an analytic controller, it also improves a contact-rich insertion task. Code and trained policies will be released.
Comments13 pages, 8 Figures, 2 Tables