arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06465eess.AS

当层选择误导语音抑郁检测

When Layer Selection Misleads Speech Depression Detection

Paula A. Perez-Toro, David Gimeno-Gómez, Daniel Rückert, Andreas Maier

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示在抑郁检测中,于评估数据上选择最佳编码器层会虚增AUC最多0.09,并证明该偏差源于选择而非信号,提出嵌套模型选择协议应作为标准,且发现紧凑的情感-韵律模型在公平比较下具有竞争力。

中文摘要 AI 辅助

预训练语音表征越来越多地用于抑郁检测,但在用于评估的同一数据上选择最佳编码器层会偏置所报告的性能。在DAIC-WOZ上,通过对五个深度编码器家族进行重复嵌套交叉验证,朴素的25层最佳探测将AUC最多虚增0.09。在基于真实潜在表征的 shuffled 标签上,朴素的25层最佳探测仍达到AUC 0.59,而嵌套选择下为0.50,且该偏差随探测层数增加和样本量减小而增大,从而将选择而非信号确认为原因。该效应在第二个临床语料库和语言(Androids,意大利语)上得到复现。在无泄漏选择下,跨表征家族的公平比较确定了一个紧凑、任务对齐的情感-韵律模型,其与显著更复杂的SSL、深度学习、音频LLM和基于编解码器的方法相比具有竞争力,同时使用的可训练下游参数约少10^3倍。将分析扩展到个体PHQ-8症状表明,相同的选择偏差在此更细粒度层面上持续存在。我们主张,在小型临床语音数据集上探测编码器层表征时,嵌套模型选择协议应成为标准做法。

英文摘要

Pretrained speech representations are increasingly used for depression detection, but selecting the best encoder layer on the same data used for evaluation biases reported performance. On DAIC-WOZ, with repeated nested cross-validation across five deep encoder families, naive best-of-25-layer probing inflates AUC by up to $0.09$. On shuffled labels over the \emph{real} latent representations, naive best-of-25 still reaches AUC $0.59$ versus $0.50$ under nested selection, and this bias grows with the number of probed layers and with smaller samples, isolating selection, not signal, as the cause. The effect replicates on a second clinical corpus and language (Androids, Italian). Under leakage-free selection, a fair comparison across representation families identifies a compact, task-aligned affect--prosody model as competitive with substantially more complex SSL, deep-learning, audio-LLM, and codec-based approaches, while using $\sim\!10^{3}\times$ fewer trainable downstream parameters. Extending the analysis to individual PHQ-8 symptoms shows that the same selection bias persists at this finer-grained level. We argue nested model-selection protocols should be standard when probing encoder layer representations on small clinical speech datasets.

补充信息

↑