arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于唇腭裂语音自动语音识别(ASR)的边缘AI实现的公平性评估

Fairness Evaluation of Edge-AI Implementation for Cleft Lip and Palate Speech ASR

Susmita Bhattacharjee, Himashri Deka, H. S. Shekhawat, S. R. M. Prasanna

arXiv 2609.03982首次发表:更新:

AI 中文总结

本研究提出感知严重程度的Whisper-small微调ASR框架,在边缘设备实现低延迟唇腭裂语音识别,降低跨严重程度组的性能差异,提升人机交互可访问性。

AI 中文摘要

自动语音识别(ASR)对唇腭裂(CLP)患者而言仍具挑战性,原因在于病理语音数据有限,且不同说话人及严重程度水平间的语音特征存在较大差异。这些识别困难会降低基于语音的人机交互的可访问性,尤其是在基于云的ASR服务不可用或不可靠时。本研究探究了一种感知严重程度且可边缘部署的ASR框架,用于使用Whisper-small模型改进唇腭裂语音的识别。该模型使用正常语音和代表轻度、中度及重度状况的唇腭裂语音的不同组合,以及仅唇腭裂的训练配置进行微调,以检验纳入不同严重程度水平如何影响识别性能和跨说话人的公平性。预训练模型的合并词错误率(WER)和音素错误率(PER)分别为62.46%和52.72%。感知严重程度的微调大幅提升了性能,将最佳合并WER降至22.72%,最佳合并PER降至18.44%。使用更广泛的唇腭裂严重程度水平表示进行训练,也在识别准确率和跨严重程度组的性能一致性之间提供了最佳整体平衡。在NVIDIA Jetson平台上的部署表明,所有微调模型均可实现实时推理,实时因子为0.167-0.171,峰值GPU内存使用量约为566 MB。结果表明,在ASR适配过程中纳入严重程度多样性可大幅改进唇腭裂语音的识别,同时减少跨严重程度组的性能差异。所提出的方法进一步支持在边缘设备上实现低延迟、不依赖互联网的语音交互,为唇腭裂患者提供更具可访问性和包容性的基于语音的人机交互。

英文摘要

Automatic speech recognition (ASR) remains challenging for individuals with cleft lip and palate (CLP) because of limited pathological speech data and large variations in speech characteristics across speakers and severity levels. These recognition difficulties can reduce the accessibility of voice-based human-computer interaction, particularly when cloud-based ASR services are unavailable or unreliable. This work investigates a severity-aware and edge-deployable ASR framework for improving recognition of CLP speech using Whisper-small. The model was fine-tuned using different combinations of normal and CLP speech representing mild, moderate, and severe conditions, together with a CLP-only training configuration, to examine how the inclusion of different severity levels influences recognition performance and fairness across speakers. The pretrained model produced pooled word error rate (WER) and phoneme error rate (PER) values of 62.46% and 52.72%, respectively. Severity-aware fine-tuning substantially improved performance, reducing the best pooled WER to 22.72% and the best pooled PER to 18.44%. Training with a broader representation of CLP severity levels also provided the best overall balance between recognition accuracy and performance consistency across severity groups. Deployment on an NVIDIA Jetson platform demonstrated real-time inference for all fine-tuned models, with real-time factors of 0.167-0.171 and peak GPU memory usage of approximately 566 MB. The results demonstrate that incorporating severity diversity during ASR adaptation can substantially improve recognition of CLP speech while reducing performance disparities across severity groups. The proposed approach further enables low-latency, Internet-independent speech interaction on edge devices, supporting more accessible and inclusive voice-based human-computer interaction for individuals with CLP.

Comments8 pages, 1 figure, Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑