arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30456cs.LG

用于婴儿哭声分析的自监督预文本任务:受控比较及Donateacry上的警示性结果

Self-Supervised Pretext Tasks for Infant Cry Analysis: A Controlled Comparison and a Cautionary Result on Donateacry

Luigi Simeone

首次发表
浏览论文内容

中文总结 AI 辅助

本文对比六种自监督预文本任务用于婴儿哭声分析,发现Donateacry哭声原因分类任务瓶颈在标签而非模型,且评估协议泄漏会虚高性能,有效样本量为婴儿数量。

中文摘要 AI 辅助

在固定预算(即相同的117万参数紧凑编码器、相同的115小时经许可验证的公开预训练音频、所有候选模型采用相同评估协议)下,我们比较了六种用于婴儿哭声分析的自监督预文本任务。在哭声检测任务中,重构目标占主导地位,即使编码器在预训练期间从未观测到哭声,基于掩码频谱图编码器的线性探针在按受试者划分的测试集上也达到了0.988的AUC。在Donateacry(哭声原因的事实上的公开基准)上的哭声原因分类任务中,所有编码器的表现均与随机猜测相当(5类的宏AUC为0.38至0.54),且在1.8小时真实哭声上进行的域适应或端到端微调均未改变结果。由于参数规模大80倍的冻结HuBERT-base也呈现相同模式,瓶颈必然在于标签而非模型容量。随后,我们仅修改评估协议,就在自身系统上复现了Donateacry文献中90%以上的准确率:按片段划分将准确率提升至85.2%(仅略高于83.8%的多数类基线),在划分前应用数据增强则将准确率提升至97.9%,达到报告的最新水平,而同一模型在按受试者划分下的宏AUC为0.49。在无数据泄漏的划分下,将标注集扩增20倍(声码器说话人扰动与噪声混合,共21小时),跨受试者AUC保持不变:对于该任务,有效样本量是婴儿的数量。我们发布代码、随机种子及每段音频的许可清单。

英文摘要

We compare six self-supervised pretext tasks for infant cry analysis under a fixed budget, meaning the same compact encoder of 1.17M parameters, the same 115 hours of license-verified public pretraining audio, and the same evaluation protocol for every candidate. On cry detection the reconstructive objectives dominate, and a linear probe over a masked-spectrogram encoder reaches 0.988 AUC with subject-wise splits even though the encoder never observed a cry during pretraining. On cry-reason classification over donateacry, the de facto public benchmark for cry reasons, every encoder performs at chance (0.38 to 0.54 macro AUC over 5 classes), and neither domain adaptation on 1.8 hours of real cries nor end-to-end fine-tuning moves the result. Since a frozen HuBERT-base with 80 times more parameters shows the same pattern, the bottleneck must sit in the labels and not in model capacity. We then reproduce the 90\%+ accuracies of the donateacry literature on our own system by changing nothing but the evaluation protocol: clip-wise splits raise accuracy to 85.2% (barely above the 83.8% majority-class baseline), and applying augmentation before splitting raises it to 97.9%, matching the reported state of the art, from the same model that measures 0.49 macro AUC under subject-wise splits. Under leakage-free splits, a twentyfold augmentation of the labeled set (vocoder speaker perturbation and noise mixing, 21 hours) leaves cross-subject AUC unchanged: for this task the effective sample size is the number of infants. We release code, seeds and per-clip license manifests.

补充信息

↑