基于Normal-Anchored一阶模型无关元学习的Whisper微调,用于提升唇腭裂语音识别的公平性
Normal-Anchored First-Order Model-Agnostic Meta-Learning based Whisper Fine-Tuning for Enhancing Fairness of Cleft Lip and Palate Speech Recognition
另 1 家 · 查看机构详情
- Indian Institute of Technology Guwahati(印度古瓦哈蒂印度理工学院)
- University of Eastern Finland(东芬兰大学)
- Brigham and Women's Hospital, Harvard Medical School(哈佛医学院布莱根妇女医院)
- Indian Institute of Information Technology Dharwad (IIIT Dharwad)(达尔瓦德印度信息技术学院(IIIT达尔瓦德))
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出NA-FOMAML微调Whisper以提升唇腭裂语音识别公平性,在两个数据集上验证了全编码器微调策略的效果,发现重度语音仍需针对性优化。
中文摘要 AI 辅助
针对唇腭裂(CLP)语音的自动语音识别(ASR)存在困难,因为其声学和发音模式随严重程度不同而变化。这种变异性会降低预训练ASR系统的性能,而传统微调在低资源、异质CLP条件下可能无法很好地泛化。本研究提出Normal-Anchored一阶模型无关元学习(NA-FOMAML),用于将Whisper适配到CLP语音。该方法采用一阶双层元学习框架,其中正常语音在内部循环中作为稳定的支撑条件,而CLP严重程度组在外部循环中用于提升适配后的鲁棒性,此设计旨在缩小正常语音与病理语音之间的性能差距。实验在NMCPC和AIISH数据集上开展,采用四种正常锚定训练配置,评估了冻结编码器、全编码器及选定Whisper编码器层的微调策略,包括层0-5、4-11、6-11、8-11,同时适配解码器和投影头。结果显示,仅用正常语音进行外部循环训练是不足的:在NMCPC数据集上,正常到正常+轻度+中度的全编码器微调,正常、轻度、中度、重度语音的词错误率(WER)分别为4.40%、5.53%、16.14%、52.07%;在AIISH数据集上,正常到正常+轻度+中度+重度的全编码器微调,对应WER分别为2.48%、19.66%、14.05%、57.50%。基于转录的音素类别分析表明,重度CLP语音在擦音、塞擦音、鼻音、流音、塞音及元音上均存在高错误率。总体而言,NA-FOMAML提升了跨严重程度的鲁棒性,但重度语音仍需严重程度感知采样、音素感知损失函数及针对压力辅音和共振相关失真的数据增强。
英文摘要
Automatic speech recognition (ASR) for cleft lip and palate (CLP) speech is difficult because acoustic and articulatory patterns vary across severity levels. This variability reduces the performance of pretrained ASR systems, and conventional fine-tuning may not generalize well under low-resource, heterogeneous CLP conditions. This work proposes Normal-Anchored First-Order Model-Agnostic Meta-Learning (NA-FOMAML) for adapting Whisper to CLP speech. The method uses a first-order bilevel meta-learning framework in which normal speech is used in the inner loop as a stable support condition, while CLP severity groups are used in the outer loop to improve post-adaptation robustness. This design aims to reduce the performance gap between normal and pathological speech. Experiments are conducted on the NMCPC and AIISH datasets using four normal-anchored training configurations. Frozen encoder, full encoder, and selected Whisper encoder-layer tuning strategies are evaluated, including layers 0--5, 4--11, 6--11, and 8--11, with decoder and projection-head adaptation. Results show that outer-loop training with only normal speech is insufficient. For NMCPC, full encoder tuning with Normal to Normal+Mild+Moderate gives WERs of 4.40%, 5.53%, 16.14%, and 52.07% for normal, mild, moderate, and severe speech. For AIISH, full encoder tuning with Normal to Normal+Mild+Moderate+Severe gives WERs of 2.48%, 19.66%, 14.05%, and 57.50%. A transcription-based phoneme-category analysis shows that severe CLP speech has high error rates across fricatives, affricates, nasals, liquids, plosives, and vowels. Overall, NA-FOMAML improves cross-severity robustness, but severe speech still requires severity-aware sampling, phoneme-aware loss functions, and augmentation targeting pressure consonant and resonance-related distortions.