一个阈值并不适用于所有语言:面向可靠高效低资源文本分类的语言条件式延迟决策
One Threshold Does Not Fit All Languages: Language-Conditional Deferral for Reliable and Efficient Low-Resource Text Classification
浏览论文内容
中文总结 AI 辅助
本研究针对低资源文本分类中单一置信度阈值无法覆盖所有语言的问题,提出按语言分别估计阈值,无需重训即可使各语言准确率达标,并揭示人工审查成本的不均衡性。
中文摘要 AI 辅助
在全球南方,即非洲、亚洲和拉丁美洲的低收入国家,那里使用着世界上大多数语言,部署的文本分类器通常运行在普通CPU上,用单一模型服务多种语言,每种语言只有少量标注样本,并依赖人工来发现错误。这样的系统只有在能够承诺其出错频率时才有用:即它自行分配的标签中,最多只能有固定比例的错误,其余所有情况都必须交由人工处理。分裂共形预测通过一个单一的置信度阈值来实现这一承诺,该阈值通常在跨语言合并的验证数据上估计。我们探究这一承诺是否对每种语言都成立,结果发现并非如此。在MasakhaNEWS(16种非洲语言)和AfriSenti(12种语言外加两种训练中从未见过的语言)上,合并阈值平均达到90%的目标,但索马里语覆盖率仅为77.5%,提格里尼亚语为83.7%,两种未见语言分别为77.5%和81.2%。为每种语言单独估计一个阈值,无需重新训练即可使每种语言达到89.1%至91.0%之间,并揭示了实现这一承诺的成本不平等:要维持承诺,意味着需将43%的索马里语新闻和超过80%的阿姆哈拉语及聪加语推文发送给人工处理,而尼日利亚皮钦语新闻则不到8%。每种语言一两百个标注样本就足够了,模型在单个CPU核心上几分钟内即可训练完成,因此这一解决方案是负担得起的:逐语言进行校准、报告并预算人工审查。
英文摘要
In the Global South, the lower-income countries of Africa, Asia, and Latin America where most of the world's languages are spoken, a deployed text classifier usually runs on ordinary CPUs, serves many languages with a single model, has few labeled examples in any of them, and relies on people to catch its mistakes. Such a system is only useful if it can promise how often it will be wrong: at most a fixed fraction of the labels it assigns on its own may be incorrect, and everything else must go to a person. Split conformal prediction delivers this promise through a single confidence threshold, normally estimated on validation data pooled across languages. We ask whether the promise reaches every language, and it does not. On MasakhaNEWS (16 African languages) and AfriSenti (12 languages plus two never seen in training), a pooled threshold meets the 90% target on average but covers Somali at 77.5%, Tigrinya at 83.7%, and the two unseen languages at 77.5% and 81.2%. Estimating one threshold per language brings every language to between 89.1% and 91.0% without retraining, and it shows how unequal the cost of the promise is: keeping it means sending 43% of Somali news and over 80% of Amharic and Xitsonga tweets to a person, against under 8% of Nigerian Pidgin news. One or two hundred labels per language are enough and the models train in minutes on one CPU core, so the fix is affordable: calibrate, report, and budget human review one language at a time.
发表机构
- University of Missouri(密苏里大学)
- Amar Bio Tech Pvt Ltd(阿玛尔生物科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。