MetaPerch:从元数据学习生物声学基础模型
MetaPerch: Learning from metadata for bioacoustics foundation models
AI总结:
研究利用生物声学数据中心的元数据,将位置、时间等元数据作为辅助监督信号,引入MetaPerch基础模型,提升物种识别性能,还对9种元数据源在17个数据集上的效果做了实证研究。
AI中文摘要:
生物声学基础模型依赖像Xeno-Canto这样的大规模公民科学平台获取地理和生态多样的数据。近期研究表明仅靠监督在这些大规模数据上训练能产生最优物种检测模型,但社区驱动的数据中心里记录的元数据仍有未利用潜力。本文探索将位置和时间等元数据用作辅助监督信号,让模型利用物种-元数据相关性。辅助元数据损失提供额外信息,促使模型有更丰富、稳健的表示,更好应对现实中被动声学监测的挑战。我们引入MetaPerch,它在多领域有强物种识别性能,并对9种元数据源在17个生物声学数据集上的效果进行了实证研究。
英文摘要:
Bioacoustic foundation models rely on large-scale citizen science platforms like Xeno-Canto for geographically and ecologically diverse data. Recent work has shown that supervision alone can produce SotA species detection models when trained on this large-scale data -- however, there remains unutilized potential in the form of recording metadata readily available within these community-driven data hubs. In this work, we explore the use of metadata -- such as location and time -- as auxiliary supervision signals, allowing the model to leverage species-metadata correlations in its learned representation. Auxiliary metadata losses provide additional information beyond vocalizations alone that can encourage a richer, more robust representation that generalizes better to species distribution and acoustic domain shifts -- important challenges for deployment in real-world passive acoustic monitoring (PAM) settings. We introduce MetaPerch, a new foundation model that achieves strong species identification performance across multiple challenging domains and present an extensive empirical study of the effects of 9 diverse metadata sources on 17 bioacoustic datasets.