发表机构
Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对BirdCLEF+ 2026中动物发声多标签检测任务,先构建监督基线,后对比神经音频编解码器的编解码表示与基础嵌入的语义表示,比较两个生物声学专家模型和四个基于令牌的编码器,探究基于令牌的表示能否竞争。
AI 中文摘要
本文详细介绍了DS@GT ARC团队针对BirdCLEF+ 2026的方法,即对潘塔纳尔湿地声景中的动物发声进行多标签检测。2026年版本增加了约一小时的标记声景,使任务转向适合标记集的监督管道。首先构建了一个有竞争力的监督基线,在90分钟CPU预算内,在排名1894时达到了0.936的私有排行榜分数。其次,对比神经音频编解码器的编解码表示与基础嵌入的语义表示,询问基于令牌的表示是否能竞争。比较了两个生物声学专家模型和四个在AudioSet上训练的基于令牌的编码器。
英文摘要
This paper details the DS@GT ARC team's approach to BirdCLEF+ 2026, multi-label detection of animal vocalizations in soundscapes from the Pantanal wetlands. The 2026 edition adds about an hour of labeled soundscapes, shifting the task toward supervised pipelines fit to the labeled set. First, we build a competitive supervised baseline that ensembles a frozen Perch v2 backbone, a trained HGNetV2-B0 sound-event-detection network, and a non-bird prototypical head, reaching a private leaderboard score of 0.936 at rank 1894 within a 90-minute CPU budget. Second, we ask whether token-based representations can compete, contrasting codec representations from neural audio codecs against semantic representations from foundational embeddings. We compare two bioacoustic specialist models against four token-based encoders trained on AudioSet. The repository for this work can be found at https://github.com/dsgt-arc/birdclef-2026.