Benchmarking Commercial Speech Recognition and Multimodal Large Language Models on Dysarthric Speech: Severity-Stratified Baselines and Architecture-Specific Prompting Effects
基于商业自动语音识别和多模态大语言模型的口吃语音零样本识别
专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);prompting(title,comments)
AI总结 本研究评估了商业ASR和MLLM在口吃语音识别中的性能,发现GPT-4o在重度口吃情况下显著降低WER,而Gemini变体表现下降,为辅助语音接口技术选择提供实证依据。
Comments 32 pages. Revised peer-reviewed version. Updated dataset filtering and transcript normalization, expanded four-condition prompting ablation and error analysis, and revised statistical reporting. Published in International Journal of Intelligent Systems