ME-Dex 1.0:将异构触觉感知引入世界动作建模
ME-Dex 1.0: Bringing Heterogeneous Tactile Sensing into World Action Modeling
浏览论文内容
中文总结 AI 辅助
提出ME-Dex-1.0,一个统一的世界动作触觉模型,通过Mixture-of-Transformers架构联合建模视觉、触觉和动作,并利用智能体数据引擎解决数据稀缺问题,提升机器人操作性能。
中文摘要 AI 辅助
世界动作模型将视频模型的预测能力引入机器人动作生成,为建模未来视觉状态提供了丰富的基础。触觉感知通过物理交互的直接测量补充了这一基础。一些现有方法使用触觉特征作为条件输入,而不联合预测未来的触觉状态、视觉观察和动作。我们的关键见解是,触觉信号与视频一样,提供了对不断演变的世界状态的观察,应作为未来观察与视频一起建模。我们提出了ME-Dex-1.0(MachEmbodied-Dex-1.0),一个统一的世界动作触觉模型,用于联合视觉、触觉和动作学习。ME-Dex-1.0采用Mixture-of-Transformers架构,包含一个视频专家、一个触觉专家和一个动作专家,均使用流匹配进行训练。我们使用共享注意力在中间层连接专家,使得在联合去噪过程中,动作生成能够利用视觉和触觉动态的学习表示。为了支持多源异构触觉输入,一个规范手模型和一个统一触觉自编码器将来自不同具身和传感布局的触觉观察映射到共享的空间和潜在空间。为了解决配对的视觉、触觉和动作数据可用性有限的问题,我们开发了智能体触觉数据引擎,一个基于智能体的数据生产平台。它通过模拟中轨迹重放期间从力传感器直接记录的触觉数据,补充了RoboTwin和DexJoCo。在RoboTwin、DexJoCo和ManiFeel模拟平台上的实验,以及真实机器人评估,展示了使用配备触觉传感的夹爪和灵巧手时操作性能的提升。
英文摘要
World Action Models bring the predictive capabilities of video models into robot action generation, providing a rich foundation for modeling future visual states. Tactile sensing complements this foundation with direct measurements of physical interaction. Some existing methods use tactile features as conditioning inputs without jointly predicting future tactile states, visual observations, and actions. Our key insight is that tactile signals, like video, provide observations of the evolving world state and should be modeled as future observations alongside video. We present ME-Dex-1.0 (MachEmbodied-Dex-1.0), a unified World Action Tactile Model for joint visual, tactile, and action learning. ME-Dex-1.0 adopts a Mixture-of-Transformers architecture comprising a Video Expert, a Tactile Expert, and an Action Expert, all trained with flow matching. We use shared attention connects the experts in intermediate layers, allowing action generation to draw on learned representations of visual and tactile dynamics during joint denoising. To support multi-source heterogeneous tactile inputs, a Canonical Hand Model and a Unified Tactile Autoencoder map tactile observations from different embodiments and sensing layouts into shared spatial and latent spaces. To address the limited availability of paired visual, tactile, and action data, we develop the Agentic Tactile Data Engine, an agent-based data production platform. It supplements RoboTwin and DexJoCo with tactile data recorded directly from force sensors during trajectory replay in simulation. Experiments on the RoboTwin, DexJoCo, and ManiFeel simulation platforms, together with real robot evaluations, demonstrate improved manipulation performance using both grippers and dexterous hands equipped with tactile sensing.
发表机构
- Li Auto Inc(理想汽车)
机构由 AI 辅助整理,请以论文原文为准。