Multi-Modal World Model for Physical Robot Interactions: Simultaneous Visual and Tactile Predictions for Enhanced Accuracy
多模态世界模型用于物理机器人交互:同时进行视觉和触觉预测以提高准确性
机构 * University of Lincoln(林肯大学) ; University of Sheffield(谢菲尔德大学)
专题命中 视频多模态 :multi-modal(title);分类 cs.CV、cs.AI
AI总结 本文提出多模态世界模型,通过整合视觉和触觉信息提升机器人物理交互的预测精度,特别是在物理模糊场景中表现更优。
Comments This paper is accepted for publication in Robotics and Autonomous Systems