arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29760cs.RO

PolyUMI:面向物体推断与操作的可及视觉-触觉-音频数据采集

PolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation

  • TU Darmstadt(达姆施塔特工业大学)
  • Robotics Institute Germany(德国机器人研究所)

机构由 AI 辅助整理,请以论文原文为准。

Conor W. Hayes, Rickmer Krohn, Aravind Ramaswami, Anunth Ramaswami, Nils Dengler, Kevin M. Lynch, J. Edward Colgate, Georgia Chalvatzaki, Matthew L. Elwin

AI总结:

PolyUMI与VisTA提供开源可及的多模态数据采集与策略学习管道,通过视觉-触觉-音频集成提升物体推断与操作性能。

AI中文摘要:

人类通常依赖视觉、触觉、听觉和本体感觉来感知接触并在操作过程中调整动作。因此,为机器人提供类似的响应能力需要能够保留和使用这些互补感官信号的硬件。然而,大多数模仿学习系统主要通过视觉和本体感觉观察示范,限制了对难以通过视觉推断的接触信息的获取。我们提出了PolyUMI,一个用于可扩展的视觉-触觉-音频示范采集和机器人部署的开源平台。其轻量级无线手持夹持器记录同步的腕部相机、光学触觉、接触音频和本体感觉观测,无需有线工作站。同一传感手指可转移到机器人末端执行器,保持示范采集与策略执行之间的传感几何一致性。为有效利用这些异构观测,我们进一步引入了VisTA,一种令牌级多模态策略,集成跨传感器和时间的信息以预测接触感知的机器人动作。涵盖物体推断、滑移控制和接触丰富操作的实验表明,触觉和音频揭示了超越视觉的任务相关信息,且VisTA与现有多模态策略相比具有竞争力或更优表现。总之,PolyUMI和VisTA提供了一个可及的管道,用于采集多模态示范并学习感知超越视觉的物理交互的策略。项目页面:此https URL

英文摘要:

Humans typically rely on vision, touch, hearing, and proprioception to perceive contact and adapt their actions during manipulation. Providing robots with comparable responsiveness therefore requires hardware that can retain and use these complementary sensory signals. Most imitation-learning systems, however, observe demonstrations primarily through vision and proprioception, limiting access to contact information that is difficult to infer visually. We present PolyUMI, an open-source platform for scalable visual--tactile--audio demonstration collection and robot deployment. Its lightweight, wireless handheld gripper records synchronized wrist-camera, optical tactile, contact-audio, and proprioceptive observations without requiring a tethered workstation. The same sensing finger can be transferred to the robot end effector, preserving the sensing geometry between demonstration collection and policy execution. To effectively use these heterogeneous observations, we further introduce VisTA, a token-level multimodal policy that integrates information across sensors and time to predict contact-aware robot actions. Experiments spanning object inference, slip control, and contact-rich manipulation show that touch and audio reveal task-relevant information beyond vision and that VisTA is competitive with or outperforms existing multimodal policies. Together, PolyUMI and VisTA provide an accessible pipeline for collecting multimodal demonstrations and learning policies that perceive physical interaction beyond vision. Project Page: https://polyumi-vista.github.io

补充信息

↑