发表机构
University of Trento; Fondazione Bruno Kessler(特伦托大学; 布鲁诺·凯斯勒基金会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出采样频率无关前端与傅里叶神经算子骨干,直接处理原生采样率录音,避免重采样和信息丢失,在84类60采样率语料上优于固定采样率基线,准确率0.906。
AI 中文摘要
传统生物声学分类模型依赖固定采样率的频谱表示,要求在不同采样率下采集的录音在分析前进行重采样。我们提出了一种采样频率无关(SFI)前端,直接以原始采样率处理每条录音,并配备具有渐进式时间尺度融合的傅里叶神经算子(FNO)骨干网络。该框架避免了固定采样率重采样和高频信息丢失,同时在不同采样率下产生固定大小的表示。训练时温和的采样率(sr)增强进一步提高了对未见采样率变化的鲁棒性。在包含84个类别和60种采样率的多分类语料库上评估,所提出的SFI-FNO配置优于固定采样率和语料库最大采样率基线,实现了0.906的准确率、0.921的平衡准确率和0.899的宏F1分数。
英文摘要
Conventional bioacoustic classification models rely on fixed-rate spectral representations, requiring recordings acquired at heterogeneous sampling rates to be resampled before analysis. We propose a Sampling-Frequency-Independent (SFI) frontend that processes each recording directly at its native sampling rate, coupled with a Fourier Neural Operator (FNO) backbone featuring progressive temporal-scale fusion. This framework avoids fixed-rate resampling and high-frequency information loss while producing fixed-size representations across sampling rates. Mild training-time sampling-rate (\textit{sr}) augmentation further improves robustness to unseen rate variations. Evaluated on a multi-taxa corpus comprising 84 classes and 60 sampling rates, the proposed SFI-FNO configuration outperforms fixed-rate and corpus-maximum-rate baselines, achieving .906 accuracy, .921 balanced accuracy, and a Macro-F1 score of .899.