arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于大象叫声分类的无参数小样本评估方法

A Parameter-Free Few-Shot Evaluation for Elephant Vocalisation Classification

Christiaan M. Geldenhuys, Thomas R. Niesler

arXiv 2608.14824首次发表:更新:

发表机构

University of Stellenbosch; Department of Electrical and Electronic Engineering(斯泰伦博斯大学; 电气与电子工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出无参数最近中心分类的小样本评估方法,在EV、LDC数据集上对比多种嵌入,发现标注示例少时该方法性能优于部分基线,示例充足时训练基线仍占优。

AI 中文摘要

我们针对大象叫声,在固定预训练声学嵌入上提出了一种无参数的情节式最近中心分类评估方法,涉及Elephant Voices(EV)和语言数据联盟(LDC)两个数据集。该方法不关注在所有可用标注数据上训练出的最优分类器,而是研究当每类标注示例数量变化时,最简单分类器的性能表现。每类由其支持集嵌入的均值表示,每个查询样本根据平方欧氏距离被分配到最近的中心。我们在与训练基线相同的交叉验证协议下,以N-way k-shot方式评估该中心分类器,使用的嵌入包括Perch(版本1)、Perch(版本2)和HuBERT(base,第2层),以及梅尔频率倒谱系数(MFCC)特征。通过对100个重采样支持集进行自助法抽样,量化采样噪声。在规模较小、资源有限的EV数据集上,使用更强的Perch(版本1)和Perch(版本2)嵌入的中心分类器,在每类仅1个示例时就超过了全训练的逻辑回归分类器,在每类2个示例时超过了更强的循环分类器。在强监督端到端基线训练所用的精简叫声类型集合上,该中心分类器在每类少量示例时,平均精度(mAP)与基线匹配并随后超越基线。在标注示例充足的更大LDC数据集上,所考虑的每个k值下,训练基线均保持优势。当每类有5个示例时,使用最强嵌入Perch(版本2)的中心分类器在EV数据集上达到0.542的mAP,在LDC数据集上达到0.368的mAP。当标注示例较少且固定嵌入已包含区分叫声类型的特征时,无参数最近中心分类是更优选择。

英文摘要

We present a parameter-free episodic evaluation of nearest-centroid classification of elephant vocalisations on fixed pretrained embeddings, for the Elephant Voices (EV) and Linguistic Data Consortium (LDC) datasets. We ask not which embedding yields the best classifier trained on all labelled data, but how the simplest classifier performs as the number of exemplars per class varies. There are no learnable parameters, because each class is modelled as the mean of its support embeddings and each query is assigned to the nearest centroid under squared Euclidean distance. Evaluation covers the fixed Perch (ver. 1), Perch (ver. 2) and HuBERT (base, layer 2) embeddings, alongside mel frequency cepstral coefficient (MFCC) features, $N$-way $k$-shot, under the same stratified $K$-fold cross-validation protocol as the trained classifiers. None of these embedding models was trained to distinguish elephant call types. On the smaller EV dataset the centroid classifier is markedly data-efficient. Using Perch (ver. 1) or Perch (ver. 2) embeddings it overtakes in mean average precision (mAP) the fully-trained logistic regression (LR) baseline from one or two exemplars and the recurrent baseline from two. Over the reduced set of call types on which the strongly-supervised end-to-end baseline was trained, the centroid classifier using Perch (ver. 2) embeddings overtakes that baseline in mAP as well, from two exemplars. On the larger LDC dataset the recurrent baselines retain their advantage for all considered values of $k$. Only LR is overtaken, and only in mAP. Nearest-centroid classification is therefore preferable precisely when exemplars are few and the fixed embedding already separates the call types.

Comments10 pages, 4 figures, 2 tables. Camera-ready version accepted at SATNAC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑