arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01865cs.CL

构音障碍语音自动语音识别(ASR)中的语音条件影响分析:分层探测研究

Analyzing Speech Condition Effects in Dysarthric ASR: A Layer-wise Probing Study

Darwin Jelestin Muthu, Navya Gupta, Wei Lin Tay, Zhengchen Zhang, Daniel Wang Zhengkui, Rong Tong

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过分层探测分析揭示构音障碍语音对ASR不同层级表征的影响,发现高层受紊乱语音影响更大,基于此提出的分层LoRA适配可高效提升低资源普通话构音障碍ASR性能。

中文摘要 AI 辅助

自动语音识别(ASR)在构音障碍语音上的性能会急剧下降,但构音障碍如何改变模型内部表征的问题尚未得到充分探索。我们针对Transformer ASR编码器在普通话构音障碍语音上开展分层探测分析,设置三种转录匹配条件:原始构音障碍语音、说话人条件零样本TTS重合成、无条件TTS。探测结果揭示了任务依赖的层级结构:构音障碍语音的音素边界信息在所有层均保持较弱,音素身份可在高层恢复,最深层编码了识别难度;声调敏感评估显示普通话词汇声调是持续的错误来源。跨条件相似性随深度增加而发散,表明紊乱语音对高层表征的影响大于低层声学特征。基于这些发现,第7层的单层LoRA、子集层5-8的适配分别达到全编码器适配相对性能差距3.5%和2.48%,而高层适配对构音障碍语音效果更差。这些发现将表征分析与参数高效微调关联,为低资源普通话构音障碍ASR的层级感知适配提供了依据。

英文摘要

Automatic speech recognition (ASR) performance degrades sharply on dysarthric speech, yet how disordered articulation reshapes a model's internal representations is underexplored. We conduct a layer-wise probing analysis of a transformer ASR encoder on Mandarin dysarthric speech under three transcript-matched conditions: original dysarthric speech, speaker conditioned zero-shot TTS resynthesis, and unconditioned TTS. Probing reveals a task- and condition-dependent representation hierarchy: phoneme boundary information remains weak across all layers for dysarthric speech; phoneme identity is recoverable in deep layers for synthetic speech, but remains poor for dysarthric speech; and recognition difficulty is concentrated in the deepest layers. Furthermore, lexical tone is a persistent error source across all conditions. Guided by these insights, layer-selective LoRA shows that mid-layer adaptation (layer 7 or layers 5-8) recovers near-full encoder performance on dysarthric speech within 6.67% and 2.89% relative margins while training only 0.16% and 0.65% of adapter parameters. Conversely, upper-layer adaptation benefits synthetic speech more than dysarthric speech. These findings link representation analysis to parameter-efficient fine-tuning and motivate layer-aware adaptation for low-resource Mandarin dysarthric ASR.

发表机构

  • Singapore Institute of Technology(新加坡理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑