AI 中文总结
该研究推出Wazobia Eval基准测试,含550+样本与16类情感分类,用于评估尼日利亚皮钦语的情感理解、反讽检测与文化推理,填补相关基准缺失,为尼日利亚语言AI提供基础评估设施。
AI 中文摘要
尼日利亚皮钦语是非洲使用最广泛的语言之一,但在语言模型评估领域却严重缺乏代表性。现有基准测试主要聚焦于翻译、转录或通用情感分析,未对基于文化的语言理解这一关键方面进行测评。本文推出Wazobia Eval,这是一个用于评估尼日利亚皮钦语情感理解、反讽检测与文化推理的基准测试。该基准基于人工标注的数据集构建,包含超过550个样本,以及16类情感分类体系,旨在捕捉传统情感框架未涵盖的文化特定情感表达。Wazobia Eval提供标准化评估协议与基准任务,用于测评模型在微妙的尼日利亚语言理解任务上的性能。本文呈现了该基准的设计、标注方法、分类体系开发流程及初步试点评估结果。我们的目标是为尼日利亚语言人工智能提供基础评估基础设施,并为未来研究建立可复现的基准。该数据集可通过此https URL公开获取。
英文摘要
Nigerian Pidgin is one of Africa's most widely spoken languages, yet remains severely underrepresented in language model evaluation. Existing benchmarks primarily focus on translation, transcription, or generic sentiment analysis, leaving critical aspects of culturally grounded language understanding unmeasured. We introduce Wazobia Eval, a benchmark for evaluating Nigerian Pidgin emotion understanding, sarcasm detection, and cultural reasoning. The benchmark is built on a manually annotated dataset containing over 550 examples and a 16-category emotion taxonomy designed to capture culturally specific emotional registers that are not represented in conventional sentiment frameworks. Wazobia Eval provides standardized evaluation protocols and benchmark tasks for assessing model performance on nuanced Nigerian language understanding. We present the benchmark design, annotation methodology, taxonomy development process, and preliminary pilot evaluation results. Our goal is to provide foundational evaluation infrastructure for Nigerian language AI and establish a reproducible benchmark for future research. The dataset is publicly available at https://huggingface.co/WAZOBIALABS.
CommentsGitHub: https://github.com/steffokoye/wazobia-eval Dataset: https://huggingface.co/WAZOBIALABS