AI 中文总结
本研究系统调查五个临床语音数据集中的 Clever Hans 效应,发现仅用静音片段即可达到或超过完整音频的分类性能,揭示数据集特定混杂因素,质疑语音生物标志物的可靠性并呼吁加强方法学报告与偏差缓解。
AI 中文摘要
近期研究在 Pitt 数据集中揭示了一个显著的 Clever Hans 效应:仅使用静音音频片段,阿尔茨海默病检测的准确率就接近 100%。这引发了人们对基于语音的健康数据集中隐藏混杂因素的严重担忧。我们系统地调查了五个广泛使用的临床语音语料库中是否存在类似偏差:DAIC-WoZ(抑郁症)、TORGO(构音障碍)、Neurovoz(帕金森病)、MDVR-KCL(帕金森病)和 UCLASS(口吃)。对于每个数据集,我们比较了使用音频第一秒、静音片段和完整录音的分类性能,并评估了原始信号和去噪信号。在所有数据集中,仅使用静音的分类性能经常达到或超过完整音频的性能,这表明分类性能可能受到数据集特定混杂因素的影响,而不仅仅是与疾病相关的语音特征。这些发现对基于语音的生物标志物的可靠性和泛化性提出了质疑,并呼吁更严格的方法学报告、预处理透明度和偏差缓解措施。
英文摘要
Recent work revealed a striking Clever Hans effect in the Pitt dataset, where Alzheimer's detection achieved nearly 100% accuracy using only silent audio segments. This raises serious concerns about hidden confounding factors in speech-based health datasets. We systematically investigate whether similar biases exist across five widely used clinical speech corpora: DAIC-WoZ (depression), TORGO (dysarthria), Neurovoz (PD), MDVR-KCL (PD), and UCLASS (stuttering). For each dataset, we compare classification using the first second of audio, silent segments, and full recordings, and evaluate both raw and denoised signals. Across all datasets, silence-only classification frequently matched or exceeded full-audio performance, suggesting that classification performance may be influenced by dataset-specific confounds in addition to disorder-related speech characteristics. These findings question the reliability and generalisability of speech-based biomarkers and call for stricter methodological reporting, preprocessing transparency, and bias mitigation.
CommentsAccepted by IEEE SLT 2026