AI 中文总结
研究视觉运动模仿学习,对比全局场景嵌入等传统视觉模型,测试以对象为中心的插槽表示,发现其结构优势,如更高成功率,还通过故障分类法分析故障类型,指出遮挡是主要瓶颈,为视觉运动模仿学习提供新方法和见解。
AI 中文摘要
机器人操纵策略依赖于预训练的视觉模型,这些模型给出全局场景嵌入或密集补丁网格,两者都混合了任务相关和无关特征。以对象为中心的插槽表示是一种结构化替代方案,它将特征分组到每个对象的几个插槽中。我们在ManiSkill3 PickCube-v1上测试这种结构的效果,使用冻结编码器和留出种子评估。保持策略、目标令牌、渲染和校准不变,仅改变编码器,冻结的以对象为中心的SPOT表示(DINO ViT-B/16 + 插槽注意力)成功率达到55.0±2.9%,比密集DINO全局特征基线(32.6±1.5%)高22.4%。单独增加令牌数量并无帮助,添加明确的2D空间目标和原生分辨率渲染可将系统提升至68.7±4.2%。自动运动学故障分类法区分了空间精度(Near-Miss)故障和对象跟踪(No-Grasp)故障,该分类法也适用于更难的StackCube-v1并指出遮挡是主要瓶颈。
英文摘要
Robotic manipulation policies rely on pre-trained vision models that give either a global scene embedding or a dense patch grid. Both mix task-relevant and task-irrelevant features. Object-centric slot representations are a structured alternative: they group features into a few per-object slots. We test what this structure buys on ManiSkill3 PickCube-v1, with a frozen encoder and a held-out-seed evaluation. Holding the policy, goal token, rendering, and calibration fixed and changing only the encoder, a frozen object-centric SPOT representation (DINO ViT-B/16 + Slot Attention) reaches 55.0$\pm$2.9% success, 22.4% above a dense DINO global-feature baseline (32.6 $\pm$ 1.5%), with the same trainable policy and no encoder fine-tuning. More tokens alone do not help: a dense patch grid with 16x the tokens performs no better than the global feature. Adding an explicit 2D spatial goal and native-resolution rendering raises the full system to 68.7$\pm$4.2%, just below a privileged 3D-oracle upper bound (71.7$\pm$4.1%). An automated kinematic failure taxonomy then separates spatial-precision (Near-Miss) failures from object-tracking (No-Grasp) failures: spatial grounding reduces Near-Miss while leaving No- Grasp unchanged. The same taxonomy transfers to the harder StackCube-v1 and points to occlusion as the main bottleneck.