Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
基于商业自动语音识别和多模态大语言模型的口吃语音零样本识别
专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);分类 eess.AS
AI总结 本研究评估了商业ASR和MLLM在口吃语音识别中的性能,发现GPT-4o在重度口吃情况下显著降低WER,而Gemini变体表现下降,为辅助语音接口技术选择提供实证依据。
Journal ref International Journal of Intelligent Systems, 2026; 2026:6065038