发表机构
Menlo Research(门洛研究公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对机器人理解自身状态问题,提出开普勒编码器v0.1,通过学习查询交叉注意力层融合多模态信息,经自监督训练。实验表明其仅视觉潜空间能恢复末端执行器状态,优于其他方法,单个编码器涵盖多个机器人,冻结潜空间有多种用途。
AI 中文摘要
机器人必须了解自身身体状态,但相机只能看到部分。力和接触在单帧中几乎无痕迹,原始视觉特征读取力的准确率在我们测试的每个机器人上仅为0.10或更低。我们提出了开普勒编码器v0.1,一种以机器人为优先的多模态编码器,将机器人状态视为一种模态,通过学习查询交叉注意力层将视觉、本体感觉和力/扭矩融合到单个共享潜空间中,在LeJEPA/SIGReg目标下通过掩码跨模态预测进行自监督训练。在评估时仅视觉输入,研究其是否因融合状态训练而使仅视觉潜空间携带像素中未包含的信息。在RH20T语料库上答案是肯定的,在相机最弱的地方表现突出。在测试场景中,仅视觉潜空间能恢复末端执行器状态,尤其是力,显著优于原始冻结ViT特征和计算匹配的仅视觉控制。单个与实施无关的编码器涵盖四个机器人,数据匹配控制表明这种广度反映了实施多样性而非数据量。冻结潜空间直接有用,其跨模态预测误差可作为训练无关的无效状态监测器,扩散解码器可从潜空间重建相机帧。本报告验证了单步情况,下一步是原生速率时间融合。
英文摘要
A robot must understand the state of its own body, but a camera sees only part of it. Force and contact leave almost no trace in a single frame, and raw vision features read force at $R^2$ at or below $0.10$ on every robot we test. We present Kepler-Encoder-v0.1, a robot-first multimodal encoder that treats robot state as a modality and fuses vision, proprioception, and force/torque into a single shared latent with a learned-query cross-attention layer, trained self-supervised by masked cross-modal prediction under the LeJEPA/SIGReg objective. At evaluation only vision enters, which poses a sharp question. Does fusing state into training make the vision-only latent carry anything the pixels do not already contain? On the RH20T corpus the answer is yes, precisely where the camera is weakest. On held-out scenes, the vision-only latent recovers end-effector state, and force in particular, significantly above both raw frozen-ViT features and a compute-matched vision-only control on every sensored robot, though absolute force recovery at a single timestep is modest; on motor state, which the camera largely sees, it is statistically tied with the strongest vision baselines, and it is the only feature whose latent geometry tracks state. A single embodiment-agnostic encoder covers four robots, and a data-matched control shows this breadth reflects embodiment diversity rather than data volume. The frozen latent is directly useful. Its own cross-modal prediction error is a training-free invalid-state monitor (AUROC $0.90$ on out-of-range states, $0.69$ on scene-swapped states), and a diffusion decoder (PixNerd) reconstructs the camera frame from the latent, confirming the spatial compression preserves world-state. This report validates the single-timestep case; native-rate temporal fusion is the next step.
Comments33 pages