发表机构
Institut de Biologie de l’École normale supérieure, CNRS, INSERM, Université PSL; Earth Species Project; Not Diamond; Institut du Cerveau; Sapienza University of Rome; École Normale Supérieure; Champalimaud Foundation(巴黎高等师范大学生物研究所(CNRS、INSERM、PSL大学); 地球物种项目; Not Diamond; 大脑研究所; 罗马大学; 巴黎高等师范学院; 尚帕利莫基金会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有海豚发声数据集规模小且不公开的问题,我们发布了 OpenWhistle,一个含约18万哨声、114小时录音及专家标注和评估协议的大规模公开数据集,并预训练了适应海豚声学的 Wav2Vec2.0 模型,其性能优于通用生物声学模型,为海豚交流研究奠定基础。
AI 中文摘要
近期生物声学领域的进展得益于大规模语料库和标准化基准,然而现有资源 overwhelmingly 以鸟类为主,且每个物种的数据深度较浅,限制了它们在研究单一物种交流系统结构方面的用途。这一缺口对鲸类尤为突出:尽管宽吻海豚(Tursiops truncatus)是非人类哺乳动物中复杂声音交流的一个引人注目的案例,但现有的海豚数据集规模小、碎片化且大多不公开。我们推出 OpenWhistle,这是最大的公开海豚发声数据集。它包含约 180,000 个哨声(114 小时),是在五年内从一个半自然环境中稳定的五只个体组成的海豚群中记录的,并配有一个由 8,354 个专家标注哨声组成的精选子集,以及用于哨声类型检测和分类的可复现评估协议。我们进一步发布了用于哨声检测、分割和分类的完整处理流程。为展示其实用性,我们在 OpenWhistle 语料库上预训练了一个适应海豚声学的 Wav2Vec2.0 模型,并表明它学习了有效的表征,在两项任务上均优于 AVES 和 BioLingual 等通用生物声学模型,同时为未来工作留下了有意义的提升空间。通过发布数据集、处理流程和评估协议,我们提供了第一个为训练自监督模型而量身定制的开放海豚哨声数据集,为推进海豚交流研究和开发捕捉物种内细粒度声学结构的模型奠定了基础。
英文摘要
Recent advances in bioacoustics have been driven by large-scale corpora and standardized benchmarks, yet existing resources are overwhelmingly bird-centric and shallow per species, limiting their use for studying the structure of a single species' communication system. This gap is particularly acute for cetaceans: despite bottlenose dolphins (Tursiops truncatus) being a compelling case of complex vocal communication among non-human mammals, existing dolphin datasets are small, fragmented, and largely closed. We introduce OpenWhistle, the largest publicly available dataset of dolphin vocalizations. It comprises approximately 180,000 whistles (114 hours) recorded over five years from a stable pod of five individuals in a semi-natural environment, paired with a curated subset of 8,354 expert-annotated whistles and reproducible evaluation protocols for whistle-type detection and classification. We further release the full processing pipeline for whistle detection, segmentation, and categorization. To demonstrate its utility, we pretrain a Wav2Vec2.0 model adapted to dolphin acoustics on the OpenWhistle corpus and show that it learns effective representations, outperforming general-purpose bioacoustic models such as AVES and BioLingual on both tasks while leaving meaningful headroom for future work. By releasing the dataset, pipeline, and evaluation protocol, we provide the first open dolphin whistle dataset tailored for training self-supervised models, laying the groundwork for advancing dolphin communication research and developing models that capture fine-grained acoustic structure within species.
CommentsAccepted as a Spotlight at the NeurIPS 2026 Datasets & Evaluations Track