arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15221cs.SDcs.LG

MAST:通过掩码音频预训练和自训练实现生物多样性监测的标签高效、鲁棒且可泛化的声音检测

MAST: Label-Efficient, Robust, and Generalizable Sound Detection for Biodiversity Monitoring via Masked Audio Pretraining and Self-Training

发表机构威斯康星大学麦迪逊分校
查看机构详情
  • University of Wisconsin--Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

Tianyi Xu, Daniel Pimentel-Alarcón, Zuzana Buřivalová, Claudia Solís-Lemus

首次发表
浏览论文内容

中文总结 AI 辅助

提出MAST框架,结合掩码音频预训练、对比学习和自训练,在无需额外标注下提升跨域声学检测的鲁棒性,在雨林和鸟类数据集上分别实现+0.22/+0.24和+0.12/+0.10的mAP/F1提升。

中文摘要 AI 辅助

被动声学监测可以在更大尺度上测量生物多样性,但动物发声的时频标注成本高昂、具有地点特异性,且难以在规模上持续进行。我们提出了一种标签高效的声音检测框架,该框架将掩码音频预训练与基于梅尔频谱图的轻量级检测器相结合,并通过在无标签音频上进行迭代自训练进一步提高鲁棒性。我们首先通过掩码重建在无标签录音上预训练一个基于ViT的编码器,并将该编码器迁移为检测骨干网络。为了更好地将动物声音与混杂背景分离,我们添加了一个框级对比损失,将匹配的事件区域拉近,同时将噪声负样本推远。然后,我们应用一个两阶段伪标签课程,以利用大规模无标签池而无需额外标注。我们在两个生态学上不同的领域评估性能:热带雨林声景(印度尼西亚)和地中海栖息地的鸟类发声(西班牙)。在这两个领域,掩码音频预训练和对比学习在时间和跨地点分布偏移下持续改善时频检测,而自训练在分布外性能上带来进一步提升。在雨林领域,MAST结合自训练在跨地点偏移下比最强基线实现了+0.22 mAP和+0.24 F1的提升。在鸟类领域,自训练在跨地点偏移下比最强基线实现了+0.12 mAP和+0.10 F1的提升。总体而言,我们的结果表明,MAST能够有效地将自监督音频表示从片段级任务扩展到跨多种生物声学场景的鲁棒框级定位,为有限标签下的生物多样性监测提供了一条实用路径。

英文摘要

Passive acoustic monitoring can measure biodiversity at larger scales, but time--frequency annotation of animal vocalizations is expensive, site-specific, and difficult to sustain at scale. We present a label-efficient sound detection framework that combines masked audio pretraining with a lightweight detector on mel spectrograms, then further improves robustness through iterative self-training on unlabeled audio. We first pretrain a ViT-based encoder on unlabeled recordings via masked reconstruction and transfer the encoder to a detection backbone. To better separate animal sounds from confounding background, we add a box-level contrastive loss that pulls matched event regions together while pushing noisy negatives apart. We then apply a two-stage pseudo-labeling curriculum to exploit large unlabeled pools without additional annotation. We evaluate the performance on two ecologically distinct domains: tropical rainforest soundscapes (Indonesia) and bird vocalizations in Mediterranean habitats (Spain). On both domains, masked audio pretraining and contrastive learning consistently improve time--frequency detection under temporal and cross-site distribution shift, and self-training yields further gains in out-of-distribution performance. On the rainforest domain, MAST with self-training achieves +0.22 mAP and +0.24 F1 over the strongest baseline under cross-site shift. On the bird domain, self-training achieves +0.12 mAP and +0.10 F1 over the strongest baseline under cross-site shift. Overall, our results show that MAST can effectively extend self-supervised audio representations from clip-level tasks to robust box-level localization across diverse bioacoustic settings, providing a practical path for biodiversity monitoring with limited labels.

补充信息

↑