arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于生物声学去噪的训练集合成:以小鼠为例

Training Set Synthesis for Bioacoustic Denoising: A Case Study With Mice

Reyhaneh Abbasi, Peter Balazs, Vincent Lostanlen, Clara Hollomey, Dustin J. Penn, Sarah M. Zala, Nicki Holighaus

arXiv 2608.10054首次发表:更新:

AI 中文总结

针对生物声学去噪的训练集合成方法,开发了脊线引导损失的U-Net去噪模型,以小鼠超声发声为案例,提升了脊线跟踪与USV分类性能,可推广至其他生物声学信号。

AI 中文摘要

生物声学记录常受环境噪声干扰,这使得对微弱或与噪声重叠的发声信号的分析变得复杂。卷积神经网络,尤其是U-Net架构,在语音和音乐处理中表现出强大的去噪性能。然而,将其直接应用于生物声学信号受到干净训练数据稀缺的限制。为解决这一问题,我们提出一种训练集合成方法,并开发了一种在时频域预测复比率掩码的监督去噪模型。该模型利用脊线(即频率轮廓),这些脊线代表发声信号的基频以及一个或多个谐波分量。这些脊线既用于训练集的合成,也用于设计损失函数,该函数为脊线区域分配更高的权重(脊线引导损失函数)。此加权步骤有助于网络在去噪过程中更好地保留发声信号的细节。作为案例研究,我们使用家鼠的超声发声(USV)记录评估我们的方法,家鼠的超声发声在行为生物学和神经科学中被广泛研究。在实际野外记录中,与我们之前的信号处理方法相比,所提方法增强了基频和谐波分量的脊线跟踪。此外,与基于噪声记录训练的分类器相比,基于去噪数据训练的分类器在野生和驯化小鼠的样本外噪声记录上的USV分类性能得到提升。我们提出的方法还在合成测试数据上,在广泛的输入信噪比范围内,大幅提高了尺度不变信号失真比。尽管我们的研究聚焦于USV,但所提方法应广泛适用于其他具有可跟踪脊线的生物声学信号,从而实现基于脊线的训练集合成与去噪。

英文摘要

Bioacoustic recordings are often degraded by ambient noise, which complicates the analysis of weak or noise-overlapped vocalizations. Convolutional neural networks, particularly U-Net architectures, have shown a strong denoising performance in speech and music processing. However, their direct application to bioacoustic signals is limited by the scarcity of clean training data. To address this issue, we propose a training set synthesis approach and develop a supervised denoising model that predicts a complex ratio mask in the time-frequency domain. The model leverages ridges, or frequency contours, that represent the fundamental frequency together with one or more harmonic partial components of vocalizations. These ridges are used both for the synthesis of training sets and to design a loss function that assigns higher weights to the ridge regions (ridge-guided loss function). This weighting step helps the network better preserve vocalization details during denoising. As a case study, we evaluate our approach using ultrasonic vocalizations (USVs) recordings of house mice, which are widely studied in behavioral biology and neuroscience. In actual field recordings, the proposed method enhances fundamental and harmonic partial ridge tracking compared to our previous signal-processing approach. In addition, a classifier trained on denoised data improves USV classification on out-of-sample, noisy recordings from wild and domesticated mice compared to classifiers trained on noisy recordings. Our proposed method also substantially improves the scale-invariant signal-to-distortion ratio on synthetic testing data across a wide range of input signal-to-noise ratios. Although we focus on USVs, the proposed approach should be broadly applicable to other bioacoustic signals with trackable ridges, and thus enables ridgebased training set synthesis and denoising.

Comments15 pages, 5 figures

Journal refIEEE Transactions on Audio, Speech and Language Processing 34 (2026) 3802-3816

DOI:10.1109/TASLPRO.2026.3705687

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑