发表机构
Universidade Estadual de Campinas(坎皮纳斯州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对非母语语音节奏评估中对齐易错的问题,提出基于卷积神经网络和低频幅度包络的声学评估方法,在speechocean762上训练,其包络一阶导数模型与人类评分相关性最高,为L2节奏评估提供无需对齐的稳健方案。
AI 中文摘要
非母语语音的自动化口语评估必须有效评价韵律,包括语音节奏,以与人类感知保持一致。然而,常用的节奏度量依赖于音段时长,需要额外的对齐步骤,这在包含不流畅和发音错误的非母语语音中容易出错。我们提出了一种基于声学的评估方法,采用卷积神经网络直接从语音幅度包络中提取节奏特征,其动机源于低频调制与节奏感知相关联的证据。所提出的模型在speechocean762数据集上训练,执行熟练度评分回归任务,并与基于时长的模型进行比较。结果表明,使用幅度包络一阶导数的模型在流畅度和韵律维度上与人类评分相关性最高,在较不流利的说话者中,其误差显著低于使用音段时长的模型。这些发现支持声学包络特征作为L2节奏评估的稳健且无需对齐的替代方案。代码已公开。
英文摘要
Automated Speaking Assessment of non-native speech must effectively evaluate prosody, including speech rhythm, to align with human perception. However, commonly employed rhythm metrics rely on segmental duration, requiring an additional alignment step, which is error-prone in non-native speech containing disfluencies and mispronunciations. We propose an acoustics-based assessment approach that employs a convolutional neural network to extract rhythm features directly from the speech amplitude envelope, motivated by evidence linking low-frequency modulations to rhythm perception. The proposed models are trained on a proficiency score regression task using the speechocean762 dataset and compared against duration-based models. Our results show that a model using the amplitude envelope's first derivative achieves the highest correlation with human-assigned scores on the Fluency and Prosody dimensions, producing significantly lower errors than one using segment durations among less fluent speakers. The findings support acoustic envelope features as robust, alignment-free alternatives for L2 rhythm assessment. Code is released publicly.
CommentsAccepted at IEEE Spoken Language Technology 2026 (SLT)