通过分层分析理解多语言医学自动语音识别(ASR)的适配
Understanding Multilingual Medical ASR Adaptation Through Layer-Wise Analysis
- Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU)(埃尔朗根-纽伦堡弗里德里希-亚历山大大学)
- University of Sheffield(谢菲尔德大学)
- Technical University of Munich (TUM)(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过分层分析Whisper模型,对比多种适配方案,发现微调可提升多语言医学ASR性能,不同设置下最优模型有别,且英语医学微调引发编码器主导偏移,多语言延续保留适配表征空间。
AI中文摘要:
医学自动语音识别(MedASR)需要适配专业术语、有限的标注临床数据以及多语言用例。尽管Whisper等大规模预训练ASR模型具备出色的泛化能力,但除词错误率(WER)外,人们对其经医学和多语言适配后的行为仍了解不足。本文通过对Whisper模型编码器的分层分析,探究多语言医学适配如何重塑其内部表征。我们在不同规模的Whisper模型上对比了零样本解码、仅英语微调、仅德语诊断微调、两阶段EN→EN+DE延续微调以及直接EN+DE微调这几种方案。微调可显著提升MedASR性能,但最优模型取决于适配设置:Whisper-Medium在直接EN+DE训练下取得最低英语WER(7.72%)和最低EN+DE组合WER(26.30%);仅德语的Whisper-Large-v3取得最低德语WER(44.96%),不过这是基于86条单说话人训练话语的语料库内诊断结果,而非稳健泛化。对两阶段Whisper-Small轨迹的分层分析显示,英语医学微调会引发编码器的主导偏移,而多语言延续则在很大程度上保留了适配后的表征空间;各层的领域和语言信息仍具备高可恢复性,而随WER提升,线性可恢复的错误预测线索会减弱。
英文摘要:
Medical automatic speech recognition (MedASR) requires adaptation to specialised terminology, limited annotated clinical data, and multilingual use cases. Although large-scale pretrained ASR models such as Whisper achieve strong generalisation, their behaviour after medical and multilingual adaptation remains insufficiently understood beyond word error rate (WER). This paper investigates how multilingual medical adaptation reshapes the internal representations of Whisper models through layer-wise encoder analysis. We compare zero-shot decoding, English-only fine-tuning, German-only diagnostic fine-tuning, two-stage EN->EN+DE continuation, and direct EN+DE fine-tuning across Whisper model sizes. Fine-tuning substantially improves MedASR performance, but the best model depends on the adaptation setting: Whisper-Medium gives the lowest English WER (7.72%) and the lowest combined EN+DE WER under direct EN+DE training (26.30%); German-only Whisper-Large-v3 gives the lowest German WER (44.96%), but as a within-corpus diagnostic on 86 single-speaker training utterances rather than robust generalisation. Layer-wise analysis of the two-stage Whisper-Small trajectory shows that English medical fine-tuning produces the dominant encoder shift, whereas multilingual continuation largely preserves the adapted representation space. Domain and language information remain highly recoverable across layers, while linearly recoverable error-predictive cues weaken as WER improves.