构音障碍语音自动语音识别(ASR)中的语音条件影响分析:分层探测研究
Analyzing Speech Condition Effects in Dysarthric ASR: A Layer-wise Probing Study
浏览论文内容
中文总结 AI 辅助
本研究通过分层探测分析揭示构音障碍语音对ASR不同层级表征的影响,发现高层受紊乱语音影响更大,基于此提出的分层LoRA适配可高效提升低资源普通话构音障碍ASR性能。
中文摘要 AI 辅助
自动语音识别(ASR)在构音障碍语音上的性能会急剧下降,但构音障碍如何改变模型内部表征的问题尚未得到充分探索。我们针对Transformer ASR编码器在普通话构音障碍语音上开展分层探测分析,设置三种转录匹配条件:原始构音障碍语音、说话人条件零样本TTS重合成、无条件TTS。探测结果揭示了任务依赖的层级结构:构音障碍语音的音素边界信息在所有层均保持较弱,音素身份可在高层恢复,最深层编码了识别难度;声调敏感评估显示普通话词汇声调是持续的错误来源。跨条件相似性随深度增加而发散,表明紊乱语音对高层表征的影响大于低层声学特征。基于这些发现,第7层的单层LoRA、子集层5-8的适配分别达到全编码器适配相对性能差距3.5%和2.48%,而高层适配对构音障碍语音效果更差。这些发现将表征分析与参数高效微调关联,为低资源普通话构音障碍ASR的层级感知适配提供了依据。
英文摘要
Automatic speech recognition (ASR) performance degrades sharply on dysarthric speech, yet how disordered articulation reshapes a model's internal representations is underexplored. We conduct a layer-wise probing analysis of a transformer ASR encoder on Mandarin dysarthric speech under three transcript-matched conditions: original dysarthric speech, speaker conditioned zero-shot TTS resynthesis, and unconditioned TTS. Probing reveals a task- and condition-dependent representation hierarchy: phoneme boundary information remains weak across all layers for dysarthric speech; phoneme identity is recoverable in deep layers for synthetic speech, but remains poor for dysarthric speech; and recognition difficulty is concentrated in the deepest layers. Furthermore, lexical tone is a persistent error source across all conditions. Guided by these insights, layer-selective LoRA shows that mid-layer adaptation (layer 7 or layers 5-8) recovers near-full encoder performance on dysarthric speech within 6.67% and 2.89% relative margins while training only 0.16% and 0.65% of adapter parameters. Conversely, upper-layer adaptation benefits synthetic speech more than dysarthric speech. These findings link representation analysis to parameter-efficient fine-tuning and motivate layer-aware adaptation for low-resource Mandarin dysarthric ASR.
发表机构
- Singapore Institute of Technology(新加坡理工学院)
机构由 AI 辅助整理,请以论文原文为准。