发表机构
Weihenstephan-Triesdorf University of Applied Sciences(魏恩施蒂芬-特里尔多夫应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对动物声学监测录音受不受控条件影响导致的分类挑战,本文提出不确定性感知融合双流框架,在跨物种身份评估的两类数据集上均显著优于静态拼接融合。
AI 中文摘要
人工智能在动物声学监测领域具有巨大潜力,可应用于精准畜牧养殖中的福利评估、野生动物保护及生态研究,相较于人工观察,动物发声能更早、更低成本地指示健康、压力及社会状态。然而,这类场景下的录音受环境噪声、混响、重叠叫声及传感器悄无声息的老化等不受控条件影响,导致动物发声自动分类颇具挑战。当前两种主流声学表征存在互补性局限:原始波形保留时间微结构,但在限幅和混响下性能下降;对数梅尔频谱图捕捉谐波结构,但丢失相位信息且对宽带噪声敏感。为应对这些挑战,本文提出不确定性感知融合(Uncertainty-Aware Fusion, UAF),这是一个双流框架,可为每种表征估计高斯不确定性,并通过不确定性加权进行融合,该机制无需可靠性标签,即可为更可信的表征分配更大权重。在跨物种、基于身份的评估(排除训练期间见过的所有个体)中,UAF(均值池化)在17类SoundWel猪叫声基准数据集上达到59.4%的准确率、39.7%的宏F1值,在3类DogBark数据集上达到73.1%的准确率、71.5%的宏F1值,相比静态拼接融合的相对宏F1值分别提升15.7%和20.4%。对四种时间聚合策略的消融实验表明,性能提升的主要驱动因素是不确定性融合,而非动物叫声的时间特性。
英文摘要
Artificial intelligence offers substantial potential for acoustic monitoring of animals, from welfare assessment in precision livestock farming to wildlife conservation and ecological research, where vocalizations can indicate health, stress, and social states earlier and at lower cost than manual observation. However, recordings in these settings are obtained under uncontrolled conditions, including environmental noise, reverberation, overlapping calls, and sensors that degrade without notice. As a consequence, automated classification of animal vocalizations remains challenging, and the two dominant acoustic representations show complementary limitations: raw waveforms preserve temporal microstructure but degrade under clipping and reverberation, while log-Mel spectrograms capture harmonic organization but lose phase information and are sensitive to broadband noise. To address these challenges, we propose Uncertainty-Aware Fusion (UAF), a dual-stream framework that estimates Gaussian uncertainty for each representation and fuses them via uncertainty weighting. This mechanism assigns greater weight to the more confident representation with no reliability labels required. In a cross-species, identity-based evaluation excluding all individuals seen during training, UAF (mean pooling) achieves 59.4\% accuracy / 39.7\% macro F1 on the 17-class SoundWel pig vocalization benchmark and 73.1\% accuracy / 71.5\% macro F1 on the 3-class DogBark dataset, outperforming static-concatenation fusion by 15.7\% and 20.4\% relative macro F1, respectively. Ablations over four temporal aggregation strategies show that uncertainty fusion, rather than the temporal characteristics of animal calls, is the primary driver of the performance gain.