arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NABEATs:噪声感知音频表示学习

NABEATs: Noise-Aware Audio Representation Learning

Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern, Christoph Boeddeker, Julius Richter, Jonathan Le Roux

arXiv 2607.16688首次发表:更新:

AI 中文总结

该研究提出噪声感知音频自监督学习概念,介绍基于BEATs的NABEATs。针对音频SSL模型在噪声环境下性能不佳的问题,通过辅助参考噪声输入训练NABEATs,能在推理时考虑噪声特征,提升下游任务性能并泛化到未见噪声类型。

AI 中文摘要

我们提出了噪声感知音频自监督学习(SSL)的概念,其目标是在抑制不需要的噪声的同时对音频混合进行编码,并提出了基于BEATs的噪声感知BEATs(NABEATs)来实现这个框架。音频SSL模型旨在处理各种音频信号,在嘈杂条件下无法有效聚焦于与下游任务相关的目标声音,导致性能下降。为解决此问题,NABEATs通过辅助参考噪声输入从嘈杂音频信号中训练估计干净的BEATs表示。该参考噪声使模型在推理时能考虑特定噪声特征,从而在不同操作环境中实现更好的泛化。实验评估表明,NABEATs在嘈杂条件下显著提高了各种下游任务的性能,并且对未见噪声类型也有良好的泛化能力。

英文摘要

We propose the concept of noise-aware audio self-supervised learning (SSL), whose goal is to encode audio mixtures while suppressing undesired noise, and present Noise-Aware BEATs (NABEATs) as a BEATs-based realization of this framework. Audio SSL models are designed to handle a wide range of audio signals. Consequently, under noisy conditions, they cannot effectively focus on the target sounds relevant to a downstream task, resulting in degraded performance. To address this issue, NABEATs is trained to estimate clean BEATs representations from a noisy audio signal with an auxiliary reference noise input. This reference noise enables the model to account for specific noise characteristics at inference time, thereby achieving better generalization across operating environments. Our experimental evaluations demonstrate that NABEATs significantly improves performance of various downstream tasks under noisy conditions and also generalizes well to unseen noise types.

CommentsAccepted at IWAENC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑