arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

野外自我运动监督的声音定位

Supervising Sound Localization by In-the-wild Egomotion

Anna Min, Ziyang Chen, Hang Zhao, Andrew Owens

arXiv 2610.01388首次发表:更新:

发表机构

Tsinghua University; University of Michigan; Shanghai Qi Zhi Institute(清华大学; 密歇根大学; 上海期智研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出利用自我运动作为监督信号学习双耳声音定位,结合传统双耳线索,在真实世界数据集上验证了模型的有效性。

AI 中文摘要

我们提出了一种利用自我运动作为监督信号来学习双耳声音定位的方法。在视频播放过程中,随着摄像机的移动,摄像机指向声源的方向会发生变化。我们训练一个音频模型来预测与摄像机运动的视觉估计一致的声音方向,这些视觉估计是通过多视图几何的传统方法获得的。这提供了一种微弱但丰富的监督形式,我们将其与传统的双耳线索相结合。为了评估这种方法,我们提出了一个包含真实世界音视频视频及自我运动的数据集。我们证明,我们的模型能够成功地从真实世界数据中学习,并在声音定位任务上表现良好。

英文摘要

We present a method for learning binaural sound localization using egomotion as a supervisory signal. Over the course of a video, the cameras direction to a sound source will change as the camera moves. We train an audio model to predict sound directions that are consistent with visual estimates of camera motion, which we obtain using traditional methods from multi-view geometry. This provides a weak but plentiful form of supervision that we combine with traditional binaural cues. To evaluate this method, we propose a dataset of real-world audio-visual videos with egomotion. We show that our model can successfully learn from real-world data and that it performs well on sound localization tasks

CommentsCVPR 2025 Highlight (IEEE/CVF Conference on Computer Vision and Pattern Recognition)

Journal refProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑