arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10790eess.AS

使用文本转语音技术进行第二语言英语口语评估的数据增强

Data Augmentation for L2 English Speaking Assessment using TTS

  • ALTA Institute(ALTA研究所)

机构由 AI 辅助整理,请以论文原文为准。

Stefano Bannò, Penny Karanasou, Mengjie Qian, Kate M. Knill, Mark J. F. Gales

AI总结:

研究针对第二语言英语口语评估数据稀缺问题,利用文本转语音技术,通过分析说话者-文本关系,提出先将书面回答语音化,再转为语音,并研究配对策略的数据增强框架,经评估可提升评分性能。

AI中文摘要:

自动化的第二语言口语能力评估依赖大规模标注语音数据,但与丰富的书面学习者语料库相比很稀缺。利用文本转语音(TTS)和语音克隆将书面L2表达转换为合成语音是解决数据不平衡的一个方向。书面和口语L2有根本差异,因此产生适合评估的合成L2语音需要什么成问题。我们通过使用COREFL语料库系统分析说话者-文本关系来解决,在框架中先将书面回答转换为口语风格转录本,再用TTS/语音克隆模型转为语音,研究不同的说话者-文本配对策略。在语言评估任务上评估数据增强技术,结果表明按熟练程度匹配说话者和文本能产生最稳健的合成语音,语音化可减少书面与口语差距并提高评分性能。

英文摘要:

Automated assessment of second language (L2) speaking proficiency requires substantial annotated speech data, which are scarce compared to written learner corpora. We investigate whether written L2 responses can be transformed into useful synthetic speech for proficiency assessment using text-to-speech (TTS) and voice cloning. Using COREFL, a corpus of paired spoken and written responses from L2 learners of English, we systematically study two factors: how written responses should be transformed into spoken-style language ("speechification") and how synthetic voices and texts should be paired based on shared learner attributes (proficiency level, first language, both, or neither). We generate speechified responses with a large language model and synthesise them using TTS and voice cloning, then evaluate their utility for audio-based (HuBERT) and text-based (ModernBERT) proficiency grading. Results show that pairing voices and texts based on proficiency provides a principled strategy for synthetic data generation, while speechification substantially improves the match between written and spoken L2 and improves downstream grading performance. Augmenting real training data with synthetic speechified responses improves both HuBERT- and ModernBERT-based graders, demonstrating the potential of synthetic spoken data for L2 speaking proficiency assessment.

↑