发表机构
Qatar Computing Research Institute(卡塔尔计算研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究实证比较了多种ASR适配方法在儿童语音上的效果,发现直接适配会损害成人性能,而权重空间合并(如LERP和TIES)能更好地平衡儿童适配与成人保留。
AI 中文摘要
自动语音识别(ASR)系统在儿童和非母语使用者上往往表现不佳,而将成人ASR模型适配到儿童语音可能会导致成人语音遗忘。我们研究了在阿拉伯语和英语中,儿童ASR适配与成人保持之间的权衡。我们比较了全量微调、LoRA和事后权重空间合并方法,这些方法应用于编码器-解码器、编码器-CTC和基于AudioLLM的ASR系统。实验使用了阿拉伯语母语和非母语儿童语音、英语MyST儿童语音,以及来自MGB-2和LibriSpeech测试清洁集的成人基准。我们使用WER评估识别质量,并通过保留指数、儿童适配增益和适配恢复来量化适配-保留权衡。结果表明,儿童适配是必要的,尤其是对于非母语阿拉伯语和英语儿童语音,但直接适配通常会降低成人ASR性能。双语适配比特定语言适配更稳定。权重空间合并通常能改善权衡,尤其是对于编码器-CTC、Whisper和基于AudioLLM的ASR,其中LERP偏向于成人保留,而TIES恢复了更强的儿童增益。对于编码器-解码器模型,直接双语微调在原始WER上仍然是最强的。代码和模型可在以下网址获取:此https URL。
英文摘要
Automatic Speech Recognition (ASR) systems often underperform for children and non-native speakers, while adapting adult ASR models to child speech can cause adult-speech forgetting. We study child ASR adaptation with adult retention across Arabic and English. We compare full fine-tuning, LoRA, and post-hoc weight-space merging across encoder--decoder, encoder--CTC, and AudioLLM-based ASR systems. Experiments use Arabic native and non-native child speech, English MyST child speech, and adult benchmarks from MGB-2 and LibriSpeech test-clean. We evaluate recognition quality with WER and quantify the adaptation--retention trade-off using Retention Index, Child Adaptation Gain, and Adaptation Recovery. Results show that child adaptation is necessary, especially for non-native Arabic and English child speech, but direct adaptation often reduces adult ASR performance. Bilingual adaptation is more stable than language-specific adaptation. Weight-space merging often improves the trade-off, especially for encoder--CTC, Whisper, and AudioLLM-based ASR, with LERP favoring adult retention and TIES recovering stronger child gains. For the encoder--decoder model, direct bilingual fine-tuning remains strongest in raw WER.\footnote{Code, and models are available at https://github.com/qcri/Child-ASR-Adaptation.
Commentslong paper