发表机构
University of Missouri; Government Degree College; Amar Bio Tech Pvt Ltd(密苏里大学; 政府学位学院; Amar生物科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对非洲语言文本分类,本文实证研究了标注预算与跨语言池化:约400个标签可达新闻分类90%性能,小预算下池化其他语言数据显著提升,零样本迁移效果有限,并发布代码与标注指南。
AI 中文摘要
针对非洲语言的每个文本分类器都始于一个预算问题:需要多少带标签的样本,以及来自其他非洲语言的标签能否替代它们?我们针对28个语言-任务组合(16种语言的新闻主题分类(MasakhaNEWS)和12种语言的推文情感分析(AfriSenti))实证回答了这两个问题,使用了一个字符n-gram线性模型,该模型在两个CPU核心上仅需数秒即可训练,无需预训练权重和加速器。预算从25到数千个标签的单语学习曲线显示,主题分类在中位语言中约400个标签即可达到其全数据宏F1的90%,而情感分析在12种语言中的11种中在完整训练规模下仍在提升,需要数千个标签。在预算较小时,池化基准中其他语言的完整训练数据价值巨大,而在预算较大时则毫无价值:在25个目标标签时,新闻分类平均增加0.20宏F1(林加拉语最高达0.43),情感分析增加0.08,增益在800个标签时衰减至零,且在完整规模下,池化在16种语言中的9种和12种语言中的8种中反而有害。25个目标标签加上池化数据达到的效果,与大多数新闻语言中100至400个单语标签相当。一个完整的零样本迁移矩阵显示,在没有任何目标标签的情况下,迁移仅恢复了多数类预测器与同语言模型之间差距的中位数的13%(新闻)和4%(情感),例外情况由共享文字(阿姆哈拉语和提格里尼亚语)、共享词汇(英语和尼日利亚皮钦语、阿拉伯方言)或共享标签先验解释,而非语言家族。我们发布了代码,可从公开基准文件重新生成每个数字,并将结果转化为针对无GPU构建非洲语言分类器的团队的具体标注指南。
英文摘要
Every text classifier for an African language begins with a budgeting question: how many labelled examples are needed, and can labels from other African languages stand in for them? We answer both questions empirically for 28 language-task pairs, news topic classification in 16 languages (MasakhaNEWS) and tweet sentiment in 12 languages (AfriSenti), using a character n-gram linear model that trains in seconds on two CPU cores with no pretrained weights and no accelerator. Monolingual learning curves at budgets from 25 to several thousand labels show that topic classification reaches 90\% of its full-data macro-F1 with about 400 labels in the median language, while sentiment is still improving at the full training size in 11 of 12 languages and needs thousands of labels. Pooling the full training data of the other languages in the benchmark is worth a great deal at small budgets and nothing at large ones: at 25 target labels it adds 0.20 macro-F1 on average for news (up to 0.43 for Lingala) and 0.08 for sentiment, the gain decays to zero by 800 labels, and at full size pooling hurts in 9 of 16 and 8 of 12 languages. Twenty-five target labels plus pooled data match what 100 to 400 monolingual labels achieve for most news languages. A complete zero-shot transfer matrix shows that transfer without any target labels recovers a median of only 13\% (news) and 4\% (sentiment) of the gap between a majority-class predictor and the in-language model, with the exceptions explained by shared script (Amharic and Tigrinya), shared lexicon (English and Nigerian Pidgin, the Arabic dialects), or a shared label prior rather than by language family. We release code that regenerates every number from the public benchmark files and translate the results into concrete annotation guidance for teams building African-language classifiers without GPUs.