发表机构
Earth Species Project; Queen Mary University of London(地球物种项目; 伦敦大学玛丽皇后学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出生物声学音频-语言任务基准BEANS-Next与训练资源ROOTS,发现现有模型在该领域窄任务外表现有限,经ROOTS训练的模型在BEANS-Next所有任务组均有显著进展,相关资源已开源。
AI 中文摘要
生物声学与动物行为学涵盖广泛的音频理解任务,其中许多任务有望从大型音频-语言模型的最新进展中受益。然而,该领域的进展迄今为止仅在狭窄的任务集上进行评估,主要集中在以标签为中心的生物类别识别,例如物种和叫声类型分类。在本研究中,我们推出BEANS-Next,这是一个基于生物声学任务分类体系的基准,涵盖声学感知、生物类别识别、场景理解和上下文学习。利用BEANS-Next,我们表明现有模型在现有评估强调的任务家族之外表现出有限的性能,限制了它们在更广泛生物声学应用中的实用性。为支持在这一更广泛任务空间上取得进展,我们还推出ROOTS,这是一个大规模训练资源,由经过扩展整理的真实世界数据、此前未充分利用的行为和声学元数据构建而成,在标注不足的地方补充了音频衍生信息和可扩展的合成生成。我们证明,在该数据集上进行训练可在BEANS-Next的所有任务组中取得显著进展,使音频-语言模型更接近其作为生物声学与动物行为学通用助手的潜力。为加速该领域的进展,我们开源了我们的基准、数据集和数据管道。
英文摘要
Bioacoustics and ethology encompass a wide range of audio understanding tasks, many of which stand to benefit from recent advances in large audio-language models. However, progress in the field has so far been assessed on a narrow set of tasks, primarily centered on label-centric biological category recognition, such as species and call-type classification. In this work, we introduce BEANS-Next, a benchmark grounded in a taxonomy of bioacoustics tasks spanning acoustic perception, biological category recognition, scene understanding, and in-context learning. Using BEANS-Next, we show that existing models exhibit limited performance beyond the task families emphasized by existing evaluations, constraining their usefulness for broader bioacoustic applications. To support progress on this broader task space, we also introduce ROOTS, a large-scale training resource built from expanded curated real-world data and previously underused behavioral and acoustic metadata, supplemented by audio-derived information and scalable synthetic generation where labeling is insufficient. We demonstrate that training on this dataset yields substantial progress across all task groups of BEANS-Next, moving audio-language models closer to their potential as general-purpose assistants for bioacoustics and ethology. To accelerate progress in the field, we open-source our benchmark, dataset, and data pipelines.
Comments36 pages, 7 figures. Project page: https://earthspecies.github.io/beans-next/