发表机构
King Saud University; Ibb University(沙特国王大学; 伊卜大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过阿拉伯语-英语语码转换数字语用分类任务,证明领域特定预训练特征比多语言覆盖更能决定Transformer性能,MARBERT显著优于XLM-R。
AI 中文摘要
本研究强调了领域特定预训练特征(DSPP)在Transformer性能中的作用,用于对阿拉伯语-英语语码转换话语中的数字语用学进行建模。研究评估了MARBERT和XLM-R(oBERTa),并以BERT作为通用基线。这些模型在分类语码转换社交媒体话语中上下文敏感的语用功能方面进行了评估。通过Python收集了11695条独特的X帖子并用于本研究。研究采用定量和定性的NLP方法,遵循监督式流程。研究发现,MARBERT始终优于XLM-R,验证宏F1从0.39提升至0.84,验证损失从0.55降至0.19。在独立测试集上,MARBERT达到了0.96的准确率、0.83的宏精确率、0.87的宏召回率和0.85的宏F1,而XLM-R达到了0.92的测试准确率,但宏F1显著较低,仅为0.52。类别层面的性能也支持了这一结果,MARBERT在F1改进上大幅领先XLM-R,提升范围从+0.33到+0.60,展示了在建模阿拉伯语数字语用学方面的明显优势。研究得出结论,Transformer性能更多依赖于DSPP而非仅靠多语言覆盖,因为后者并不能保证在高度专业化的语用分类任务上达到最优性能。
英文摘要
This study highlights the role of domain-specific pretraining profile (DSPP) in Transformer performance for modeling digital pragmatics in Arabic-English code-switched discourse. It evaluates MARBERT and XLM-R(oBERTa), with BERT serving as a general-purpose baseline. The models were evaluated on their ability to classify context-sensitive pragmatic functions in code-switched social-media discourse. 11695 unique X posts were collected via Python and utilized for the study. The study employs a quantitative and qualitative NLP approach, following a supervised pipeline. Findings unveil that MARBERT consistently surpasses XLM-R with validation Macro F1 increasing from 0.39 to 0.84 and validation loss decreasing from 0.55 to 0.19. On an independent test set, it achieved 0.96 accuracy, 0.83 macro precision, 0.87 macro recall, and 0.85 Macro F1, while XLM-R achieved 0.92 test accuracy but a substantially lower Macro F1 of 0.52. This was also supported by class-level performance where MARBERT outperforms XLM-R considerably with F1 improvements ranging from +0.33 to +0.60, demonstrating a clear advantage in modeling Arabic digital pragmatics. The study concludes that Transformer performance depends more on DSPP than multilingual coverage alone, as the latter does not guarantee optimal performance on a highly specialized pragmatic classification task.
Comments19 pages, 2 figures, 5 tables