发表机构
University of Helsinki(赫尔辛基大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对低资源欧洲语言数据稀缺致文本难度评估模型发展受阻的问题,提出利用机器翻译进行跨语言数据增强的策略,训练基于BERT的回归模型,实验证明该方法能显著提高难度估计准确性,为相关语言提供可行方案。
AI 中文摘要
可靠的文本难度评估是有效的文本简化工作流程和个性化学习应用的前提。然而,强大的评估模型发展受到严重阻碍,关键瓶颈在于缺乏包含细粒度难度级别(如欧洲共同语言参考标准)的专家注释语料库,尤其是低资源语言。本文针对低资源欧洲语言的数据稀缺问题,提出一种跨语言数据增强策略,利用机器翻译将高资源语言的标注资源转移到目标低资源语言。训练基于BERT的回归模型预测难度分数,并研究合成翻译数据能否有效补充原生训练集。实验表明,用机器翻译语料库增强稀缺的原生数据可显著提高难度估计的准确性,为缺乏大量专家注释的语言提供了可行解决方案。
英文摘要
Reliable Text Difficulty Assessment is a prerequisite for valid text simplification workflows and personalized learning applications. However, the development of robust assessment models is severely hindered by a critical bottleneck: the scarcity of expert-annotated corpora containing fine-grained difficulty levels (e.g., CEFR), particularly for lower-resource languages. This paper addresses this data scarcity problem in the context of a low-resource European language. We propose a cross-lingual data augmentation strategy that leverages machine translation to transfer labeled resources from high-resource languages to the target low-resource language. We train BERT-based regression models to predict difficulty scores and investigate whether synthetic, translated data can effectively supplement native training sets. Our experiments demonstrate that augmenting scarce native data with machine-translated corpora significantly improves the accuracy of difficulty estimation, offering a viable solution for languages lacking extensive expert annotations.