arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

S2A2:利用声学空间信息进行操作任务的视听模仿学习

S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information

Kaneyoshi Hiratsuka, Benjamin Yen, Ryosuke Kojima

arXiv 2607.26047首次发表:更新:

发表机构

Kyoto University; RIKEN; Institute of Science Tokyo(京都大学; 理化学研究所; 东京理科大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究利用声学信息进行操作任务的模仿学习,提出多模态框架S2A2,集成视觉与声学信息,实现多种策略集成,模拟与真实机器人实验表明其对特定任务有效且适用于现实操作。

AI 中文摘要

声学信息提供了有关物体位置、材料属性以及接触或运动引起的变化的丰富线索。本文引入了一组新的用于模仿学习的声学感知操作任务,机器人必须利用听觉线索来确定操作目标。这些任务需要进行声源定位和识别以在机器人操作中进行主动探索。此外,我们提出了一种多模态模仿学习框架Spatial-Spectral Audio Action(S2A2),它将视觉特征与声学空间和声学信号信息集成用于声学感知操作任务。我们实现了将ACT、Diffusion Policy、VQ-BeT和$\pi_0$等策略集成到框架中的S2A2模型。模拟实验表明该方法对需要位置和音色的任务最有效。真实机器人实验证实了所提出的任务和框架在现实世界操作中的适用性。

英文摘要

Acoustic information provides rich cues about object location, material properties, and changes caused by contact or motion. This paper introduces a new set of acoustic-aware manipulation tasks for imitation learning, in which robots must use auditory cues to determine manipulation targets. These tasks require sound source localization and identification for active exploration in robotic manipulation. Also, we propose a multimodal imitation learning framework, Spatial-Spectral Audio Action (S2A2), that integrates visual features with acoustic spatial and acoustic signal information for the acoustic-aware manipulation tasks. We implemented S2A2 models that integrates policies such as ACT, Diffusion Policy, VQ-BeT, and $π_0$, into our framework. Simulation experiments showed that the proposed method is the most effective for tasks requiring both position and timbre. Furthermore, real-robot experiments confirm the applicability of the proposed tasks and framework to real-world manipulation.

CommentsProject page: https://azuma413.github.io/projects/s2a2

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑