Bearings:基于一阶环境立体声的自监督声场嵌入
Bearings: Self-Supervised Soundfield Embeddings from First-Order Ambisonics
浏览论文内容
中文总结 AI 辅助
Bearings是一种自监督声场嵌入框架,从一阶环境立体声学习空间表征,可附加到冻结音频编码器,实现联合声音事件检测与定位,显著提升F分数。
中文摘要 AI 辅助
最近提出的自监督音频编码器能够学习强大的通用声音场景表征,但这些表征在空间上是盲的。为了补充声音场景缺失的空间表征,我们引入了Bearings。Bearings是一个自监督框架,从无标注的一阶环境立体声中学习声场嵌入。我们预训练了一个掩码自编码器,并配有一个解码器,该解码器以来自现成的单通道音频编码器的冻结声学嵌入为条件。我们的结果表明,生成的声场嵌入形成了一个可复用的流,可以通过轻量级可训练融合头附加到冻结的声学编码器上。在声音事件定位与检测任务中,将我们的声场嵌入与声学表征拼接,提供了缺失的空间信息,并实现了联合检测与定位,将TAU-NIGENS 2021上的位置相关F分数从低于4提升到50,在STARSS23上提升到39。据我们所知,Bearings是第一个自监督声场编码器,其嵌入可以插入冻结的声学编码器而无需重新训练任一模型。
英文摘要
Recently proposed self-supervised audio encoders learn powerful general-purpose representations of sound scenes, yet they are spatially blind. To supply the missing spatial representation of sound scenes, we introduce Bearings. Bearings is a self-supervised framework that learns soundfield embeddings from unlabeled first-order Ambisonics. We pre-train a masked auto-encoder paired with a decoder conditioned on frozen acoustic embeddings from an off-the-shelf single-channel audio encoder. Our results show that the resulting soundfield embeddings form a reusable stream that can be attached to frozen acoustic encoders with a lightweight trainable fusion head. On sound event localization and detection, concatenating our soundfield embeddings with acoustic representations provides the missing spatial information and enables joint detection and localization, raising the location-dependent F-score from below 4 to 50 on TAU-NIGENS 2021 and 39 on STARSS23. To our knowledge, Bearings is the first self-supervised soundfield encoder whose embeddings plug into frozen acoustic encoders without retraining either model.
发表机构
- Radboud University(拉德堡德大学)
- Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。