发表机构
Aalto University; University of Helsinki; Walton Institute(阿尔托大学; 赫尔辛基大学; 沃尔顿研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出结合Whisper-medium与Qwen3.5-2B的CASA架构,在Speak & Improve Corpus 2025上RMSE达0.358,参数量减半,可分离语音表达与内容,还分析了声学与内容信息的贡献及性能稳定性。
AI 中文摘要
自动口语评估(ASA)研究越来越多地采用多模态语音大语言模型来评估学习者的口语表现。然而,现有研究对声学信息和内容信息如何影响预测结果,以及所得性能的稳定性分析有限。我们提出CASA,一种结合Whisper-medium(语音编码器)和Qwen3.5-2B(大语言模型)的更简单架构,在实现最先进性能的同时,能更清晰地分离语音表达与内容。在Speak & Improve Corpus 2025数据集上,CASA的均方根误差(RMSE)为0.358,优于此前最佳RMSE,且所用推理参数约为其一半。该通用架构设计用于无需结构调整即可适配其他ASA数据集,依赖3个人工设计的流畅度特征。通过消融实验和重复运行,我们分析了声学与内容信息的单独及互补贡献,检验了性能变异性,并证明了大语言模型推理在无训练内容验证方面的潜力。
英文摘要
Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies provide limited analysis of how acoustic and content information contribute to predictions and how stable the resulting performance is. We propose CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content. On the Speak & Improve Corpus 2025, CASA achieves a root mean square error (RMSE) of 0.358, improving on the previous best RMSE while using approximately half the estimated inference parameters. The general-purpose architecture is designed for adaptation to other ASA corpora without structural changes and relies on three handcrafted fluency features. Through ablations and repeated runs, we analyze the individual and complementary contributions of acoustic and content information, examine performance variability, and demonstrate the potential of large language model reasoning for training-free content validation.
CommentsTo be submitted to ICASSP 2027. Code is available at https://github.com/aalto-speech/casa