arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经表征,自然连接:从人类语音基础模型到动物叫声迁移了什么?

Neural Representations, Natural Connections: What Transfers From Human Speech Foundation Models to Animal Vocalizations?

Tomás Arias-Vergara, Christopher Hauer, Héloïse Brotier, Elmar Nöth, Andreas Maier, Lee Koren

arXiv 2610.07107首次发表:更新:

发表机构

Friedrich-Alexander-Universität Erlangen-Nürnberg; OTH Amberg-Weiden; Bar Ilan University(弗里德里希-亚历山大大学埃尔朗根-纽伦堡; 安贝格-魏登应用科学大学; 巴伊兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估15个冻结编码器在动物叫声识别上的跨物种迁移,发现语音预训练在部分任务中优于动物预训练,且迁移效果强烈依赖层深度。

AI 中文摘要

语音预训练模型在动物生物声学中展现出潜力,但控制其跨物种迁移的因素仍知之甚少。我们评估了15个冻结编码器,涵盖单语和多语语音、说话人验证、动物生物声学及通用音频,在四个物种的呼叫者识别和三个物种的叫声类型分类任务上进行测试。最强语音预训练与动物预训练表征之间的差异范围从-0.027到+0.119 UAR;在七个设置中,语音在三个设置中显著更优(p<0.01),在其余四个设置中无显著差异。受控比较显示,增加语言覆盖范围或动物领域预训练并无系统性优势,而冻结的说话人验证嵌入迁移效果较差。结果还表明,迁移强烈依赖于层:原始波形模型在早期层达到峰值,而补丁频谱图模型在深层达到峰值。

英文摘要

Speech-pretrained models have shown promise in animal bioacoustics, but the factors governing their cross-species transfer remain poorly understood. We evaluate 15 frozen encoders spanning monolingual and multilingual speech, speaker verification, animal bioacoustics, and general audio on caller identification across four species and call-type classification across three. Differences between the strongest speech- and animal-pretrained representations range from $-0.027$ to $+0.119$ UAR; speech is significantly better in three of seven settings ($p<0.01$) and not significantly different in the remaining four. Controlled comparisons show no systematic advantage from increased language coverage or animal-domain pretraining, while frozen speaker-verification embeddings transfer poorly. The results also show that transfer is strongly layer-dependent: raw-waveform models peak early, whereas patch-spectrogram models peak deeper.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑