arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用手工制作的以MFCC为主的声学生物标志物从自发语音中进行无转录本的阿尔茨海默病轻量级检测

Transcript-Free Lightweight Detection of Alzheimer's Disease from Spontaneous Speech Using Handcrafted MFCC-Dominant Acoustic Biomarkers

Rashin Gholijani Farahani, Azam Bastanfard

arXiv 2607.10168首次发表:更新:

AI 中文总结

研究旨在从自发语音中无转录本检测阿尔茨海默病,利用痴呆症银行匹兹堡语料库,提取99个手工声学-时间特征,用WebRTC VAD分离语音,通过支持向量机评估,结果显示相关线索可促进独立于说话者的AD筛查,为部署研究奠定基础。

AI 中文摘要

早期发现阿尔茨海默病(AD)仍然困难,尤其是在神经成像昂贵或依赖语言的工具不可用时。自发语音提供了一种非侵入性信号,但当前许多方法依赖转录本/自动语音识别(ASR)或计算密集型深度模型。我们利用痴呆症银行匹兹堡语料库中的176个“曲奇盗窃”录音(88例AD,88例对照),提供了一个简单的仅音频基线来检测AD。使用WebRTC语音活动检测(VAD)分离语音和非语音。提取99个手工制作的声学-时间特征,包括停顿和流畅性统计、频谱/韵律描述符以及带有{\Delta}和{\Delta}{\Delta}的MFCC摘要。使用严格的独立于说话者的GroupShuffleSplit进行评估,记录30次迭代的性能。具有RBF核的轻量级支持向量机(SVM)在各次运行中的平均AUC为0.674。例如,单次分割的AUC为0.742,准确率为0.657。还进行了探索性紧凑特征分析,利用随机森林重要性排名的前20个子集;由于选择不在训练分割内嵌套,这些结果可能过于乐观,未用于主要结论(AUC 0.719)。结果表明,无转录本的频谱-时间和流畅性相关线索可促进从原始音频中进行独立于说话者的阿尔茨海默病筛查,为面向部署的研究奠定了实际基础。

英文摘要

It is still hard to find Alzheimer's disease (AD) early, especially when neuroimaging is expensive or tools that depend on language are not available. Spontaneous speech provides a non-invasive signal; however, numerous current methodologies depend on transcripts/ASR or computationally intensive deep models. We offer a simple, audio-only baseline for detecting AD using 176 Cookie Theft recordings from the DementiaBank Pitt corpus (88 AD, 88 controls). WebRTC voice activity detection (VAD) is used to separate speech from non-speech. We take out 99 hand-crafted acoustic-temporal features, including pause and fluency statistics, spectral/prosodic descriptors, and MFCC summaries with Δ and ΔΔ. Evaluation is performed using a stringent speaker-independent GroupShuffleSplit,documenting performance across 30 iterations. A lightweight SVM with an RBF kernel gets an average AUC of 0.674 across runs. For example, a single split has an AUC of 0.742 and an accuracy of 0.657. We also present an exploratory compact-feature analysis utilizing a Top-20 subset ranked by Random Forest importance; since selection is not nested within training splits, these results may be overly optimistic and are not employed for primary conclusions (AUC 0.719). The results indicate that transcript-free spectro-temporal and fluency-related cues can facilitate speaker-independent Alzheimer's disease screening from raw audio, establishing a practical foundation for deployment-oriented research.

Comments10 pages, 5 figures, 40 references. Submitted for peer review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑