发表机构
Earth Species Project(地球物种计划)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究利用四个预训练音频模型从物种发声恢复系统发育距离,在海洋哺乳动物和鸟类中,通用基础模型能恢复信号,特定领域预训练的BirdNET等未超越,表明预训练音频嵌入可携带进化信息,特定领域预训练非必需。
AI 中文摘要
我们探究了四个大型预训练音频模型(AST、CLAP、BEATs - bio和BirdNET),利用一个它们在训练期间都未见过的下游任务:从物种发声中恢复系统发育距离。在32种海洋哺乳动物(来自沃特金斯海洋哺乳动物声音数据库的1754条录音)中,基础模型在26种鲸类动物中恢复了强烈的系统发育信号(CLAP r = 0.82,BEATs - bio r = 0.82,AST r = 0.74;所有p < 0.001),而手工制作的MFCC特征(105维)则未发现信号(r = 0.040,p = 0.338)。在20种鸟类中重复分析,通用基础模型再次恢复了信号(AST r = 0.55,CLAP r = 0.52),而BirdNET和BEATs - bio并未超越它们(r约为0.32至0.36)。预训练音频嵌入携带跨两个独立辐射的进化信息,特定领域预训练并非其出现的必要条件。
英文摘要
Do learned audio embeddings encode structure that nobody told them to encode? We probe four large pretrained audio models (AST, CLAP, BEATs-bio and BirdNET) with a downstream task none of them saw during training: recovering phylogenetic distance from species vocalizations. If the geometry of the embedding space tracks the tree of life, the representation is picking up something deeper than the labels the model was optimized for. We run Mantel tests across two independent radiations. In 32 marine mammal species (1,754 recordings from the Watkins Marine Mammal Sound Database) the foundation models recover strong phylogenetic signal within the 26 cetaceans (CLAP r=0.82, BEATs-bio r=0.82, AST r=0.74; all p<0.001), among the highest acoustic-phylogenetic correlations reported for any taxon. Hand-crafted MFCC features (105d) find nothing (r=0.040, p=0.338). The gap survives after PCA-projecting every embedding down to 105 dimensions, so it is not an artefact of representation size. It also survives a partial Mantel test controlling for dominant frequency (partial Mantel r=0.404, keeping 97% of the variance explained), so it is not just pitch in disguise. We repeat the analysis on 20 bird species using the Jetz et al. (2012) phylogeny, and this time add BirdNET, a classifier trained end-to-end on around 6,000 bird species. The general-purpose foundation models recover the signal again (AST r=0.55, CLAP r=0.52). The unexpected result is that neither BirdNET nor the bioacoustic BEATs-bio beat them (r around 0.32 to 0.36). Matching the training domain to the target taxon does not, by itself, help. Pretrained audio embeddings carry evolutionary information across two independent radiations, and domain-specific pretraining is not required for it to emerge.
Comments17 pages, 5 figures, 2 supplementary tables. Code, embeddings and derived matrices: https://github.com/rinvictor/bioacoustic-phylogeny-embeddings