发表机构
University of Toronto; North York General Hospital(多伦多大学; 北约克综合医院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对ALS检测和进展预测问题,提出基于预训练SSL嵌入构建k近邻图的个体层面图框架,比较多种SSL前端和图神经网络,最佳配置在相关任务上表现优于基线,凸显结合两者用于低资源ALS监测的潜力。
AI 中文摘要
肌萎缩侧索硬化症(ALS)逐渐损害言语运动控制,使声学分析成为用于严重程度和进展估计的有前景的生物标志物。我们提出一个个体层面的图框架,将多个发声记录聚合到一个由2秒片段的预训练自监督学习(SSL)嵌入构建的唯一k近邻图中。我们在SAND数据集任务(339名参与者:205名ALS患者、134名对照)上比较了四个SSL前端(wav2vec 2.0、HuBERT、data2vec - audio和UniSpeech - SAT)和五个图神经网络(GCN、残差GCN、GAT、GraphSAGE和GIN),任务包括5类构音障碍严重程度和4类ALSFRS - R进展预测。在官方验证集上,最佳配置(HuBERT + GIN)在任务1中实现了0.73的宏F1,在任务2中实现了0.69,优于SAND验证基线(0.61和0.58)。这些结果凸显了将图神经网络与预训练的跨语言语音表示相结合用于低资源ALS检测和进展监测的潜力。
英文摘要
Amyotrophic lateral sclerosis (ALS) progressively impairs speech motor control, making acoustic analysis a promising biomarker for severity and progression estimation. We propose a subject-level graph framework that aggregates multiple phonation recordings into a unique k-nearest-neighbor graph built from pretrained SSL embeddings of 2s segments. We compare four SSL front-ends (wav2vec 2.0, HuBERT, data2vec-audio, and UniSpeech-SAT) and five graph neural networks (GCN, residual GCN, GAT, GraphSAGE, and GIN) on the SAND dataset tasks (339 participants: 205 ALS, 134 control): 5-class dysarthria severity and 4-class ALSFRS-R progression prediction. On the official validation set, the best configuration (HuBERT+GIN) achieves macro-F$_1$ of 0.73 for Task 1 and 0.69 for Task 2, outperforming SAND validation baselines (0.61 and 0.58). These results highlight the potential of combining GNNs with pretrained cross-lingual speech representations for low-resource ALS detection and progression monitoring.
CommentsAccepted for publication at Interspeech2026