arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15666eess.AScs.SD

OpenEnded:带有人工标注和ALM监督的开放式口语能力评估语音语料库

OpenEnded: An Open-Response Speech Corpus for Speaking Proficiency Assessment with Human Annotations and ALM Supervision

Yu-Wen Chen, Eric Zhou, Evelyn Ding, Tianyi Shen, Zhou Yu, Julia Hirschberg

首次发表
浏览论文内容

中文总结 AI 辅助

针对自动口语评估数据稀缺问题,构建开放式英语口语语料库OpenEnded,提供语句级多维评分,并验证ALM伪标注及VoxPA基线模型的有效性。

中文摘要 AI 辅助

自动口语评估(ASA)的发展受到公共数据集稀缺的限制,现有大多数工作依赖于朗读语音,这限制了其在实际交流场景中的适用性。在本工作中,我们引入了OpenEnded,一个包含普通话者在开放式任务中英语练习语音的语料库。与仅提供整体熟练度评分的先前开放式数据集不同,OpenEnded提供了针对准确性、流利度和韵律的语句级评估。我们收集了约10,000条语句,并使用混合框架进行标注:其中1,000条通过多评分者评分及差异解决进行人工标注,形成高质量测试集,其余语句由音频语言模型(ALM)进行伪标注,用于训练集和开发集。我们在OpenEnded测试集上评估了ALM和现有ASA模型,并引入了VoxPA作为额外基线。结果表明,ALM生成的伪标签比原始ALM评分更能改善训练效果,而VoxPA在所有基线中取得了最佳性能。

英文摘要

The development of automated speaking assessment (ASA) is limited by the scarcity of public datasets, with most existing work relying on read-aloud speech, which limits applicability to real-world communication scenarios. In this work, we introduce OpenEnded, a corpus of English practice speech from Mandarin speakers in open-response tasks. Unlike prior open-response datasets that provide only holistic proficiency scores, OpenEnded offers utterance-level assessments of accuracy, fluency, and prosody. Approximately 10,000 utterances are collected and annotated using a hybrid framework: 1,000 are manually labeled via multi-rater scoring with discrepancy resolution to form a high-quality test set, while the remaining are pseudo-labeled by an audio language model (ALM) for training and development sets. We evaluate ALMs and existing ASA models on the OpenEnded test set and introduce VoxPA as an additional baseline. Results show that ALM-generated pseudo-labels improve training over original ALM scoring, while VoxPA achieves the best performance among all baselines.

发表机构

  • Columbia University(哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑