arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于清醒开颅术中跨说话人构音障碍检测的自监督语音表示

Self-Supervised Speech Representations for Cross-Speaker Dysarthria Detection During Awake Craniotomy

Kanthila Chinmayi, Abdallah Nassib, Misy Harrison, Panheleux Celine, Saliou Vanessa, Seizeur Romuald, Dardenne Guillaume

arXiv 2610.11825首次发表:更新:

发表机构

LaTIM UMR1101, INSERM, University of Western Brittany; University Hospital of Brest(布列塔尼西大学 LaTIM UMR1101 研究所; 布雷斯特大学医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究基于DATABRASE语料库,结合wav2vec 2.0等技术构建流程,在严格跨说话人条件下实现清醒开颅术中构音障碍检测,发现自监督语音表示对性能影响显著,为术中语言功能保护提供方法支持。

AI 中文摘要

在清醒开颅术中检测术中言语损伤对保留语言功能至关重要,但自动检测仍具挑战性,因为手术室录音存在大量声学干扰、临床相关言语事件稀少,且可用队列规模小、说话人间异质性大。本研究对一个用于区分构音障碍与正常言语的流程进行了系统性组件评估,该流程基于清醒开颅术录音的DATABRASE语料库,包含说话人 diarization(说话人分割)以隔离患者言语、结合手工声学描述符与多层wav2vec 2.0嵌入的多视图表示、基于说话人条件归一化和可迁移性的特征选择以提升跨说话人鲁棒性,以及由梯度提升第一阶段和神经网络第二阶段组成的级联分类器。评估在严格的说话人独立条件下采用留一说话人交叉验证进行。结果显示,跨说话人性能受言语表示的影响远大于分类器选择,三种分类器的AUC差异不超过4.7%,而用多层自监督表示替代传统声学描述符可使AUC提升18.2%-26.1%;受 diarization 约束的特征提取和所提分类器级联提供了额外的稳定增益。这些发现表明,在低资源术中场景中,可靠的患者特异性言语隔离和强大的预训练表示比增加分类器复杂度更重要,同时也量化了通过患者特异性术前校准可实现的潜在性能增益。

英文摘要

Detecting intra-operative speech impairment during awake craniotomy is essential for preserving language function. However, automated detection remains challenging because operating-room recordings contain substantial acoustic interference, clinically relevant speech events are rare, and available cohorts are small and heterogeneous across speakers. This study presents a systematic component-wise evaluation of a pipeline for distinguishing dysarthric from no-trouble speech in the DATABRASE corpus of awake-craniotomy recordings. The pipeline incorporates speaker diarization to isolate patient speech, a multi-view representation combining handcrafted acoustic descriptors with multilayer wav2vec 2.0 embeddings, speaker-conditional normalization and transferability-based feature selection to improve cross-speaker robustness, and a cascaded classifier comprising a gradient-boosted first stage and a neural second stage. Evaluation was conducted under strict speaker-independent conditions using leave-one-speaker-out cross-validation. The results show that cross-speaker performance is influenced more strongly by the speech representation than by classifier choice. The AUCs of three classifiers differed by no more than 4.7%, whereas replacing conventional acoustic descriptors with the multilayer self-supervised representation produced AUC improvements of 18.2%-26.1%. Diarization-conditioned feature extraction and the proposed classifier cascade provided additional consistent gains. These findings indicate that reliable patient-specific speech isolation and strong pretrained representations are more important than increased classifier complexity in low-resource intra-operative settings. They also quantify the potential performance gains that may be achieved through patient-specific preoperative calibration.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑